VnRobo
Về chúng tôiBảng giáBlogLiên hệ
🇺🇸ENĐăng nhậpDùng thử miễn phí
🇺🇸EN
VnRobo logo

Hạ tầng AI cho robot công nghiệp thế hệ mới.

Sản phẩm

  • Tính năng
  • Bảng giá
  • Kiến thức
  • Dịch vụ

Công ty

  • Về chúng tôi
  • Blog
  • Liên hệ

Pháp lý

  • Chính sách bảo mật
  • Điều khoản sử dụng

© 2026 VnRobo. Bảo lưu mọi quyền.

Được tạo với♥tại Việt Nam
VnRobo
Về chúng tôiBảng giáBlogLiên hệ
🇺🇸ENĐăng nhậpDùng thử miễn phí
🇺🇸EN
  1. Trang chủ
  2. Blog
  3. RL²-VLA: Offline RL Latent Steering Tăng +26% Success Rate VLA Tại Test Time
wholebody-vlavlaoffline-rlflow-matchingmanipulationtest-time-scalingpi0conformal-prediction

RL²-VLA: Offline RL Latent Steering Tăng +26% Success Rate VLA Tại Test Time

Hướng dẫn chi tiết dùng RL²-VLA để steer VLA manipulation policy bằng offline RL trong không gian latent, không cần retrain model gốc, đạt +26% real-robot success rate.

Nguyễn Anh Tuấn31 tháng 7, 202612 phút đọc
RL²-VLA: Offline RL Latent Steering Tăng +26% Success Rate VLA Tại Test Time

RL²-VLA: Offline RL Latent Steering Tăng +26% Success Rate VLA Tại Test Time, Không Cần Retrain

Vision-Language-Action (VLA) models như π0, OpenVLA đã đạt kết quả ấn tượng trong robot manipulation. Nhưng có một vấn đề cố hữu: khi gặp task hoặc ngữ cảnh khác với training data (out-of-distribution), success rate sụt mạnh. Giải pháp truyền thống — fine-tune lại model — tốn kém và chậm.

RL²-VLA từ MARMot Lab (NUS) đề xuất hướng đi khác hoàn toàn: thay vì retrain, ta bổ sung một steering policy nhỏ được train bằng offline RL, hoạt động trong không gian latent của VLA, và chỉ can thiệp khi model gốc được dự đoán sẽ thất bại. Kết quả: +26.7% success rate trên robot thực (PiperX), +14.7% trên SIMPLER benchmark, không sửa một trọng số nào trong VLA gốc.

  • Paper: RL²-VLA (arXiv 2607.26991)
  • Code: github.com/marmotlab/RL2-VLA
  • Project page: rl2-vla.github.io
  • Models: huggingface.co/rl2-vla

Khuyến nghị công cụ

Stack train/deploy cho VLA

Train trên cloud/workstation, deploy bản tối ưu xuống Jetson hoặc robot computer.

Cloud GPU for VLA / policy training Dùng cho imitation learning, diffusion policy, RL và fine-tuning model robotics. Xem cloud GPU → NVIDIA Jetson Orin NX / Orin Nano Máy deploy edge cho perception, logging và inference đã tối ưu. Xem Jetson → Hugging Face / robotics dataset hosting Lưu dataset, checkpoint và model card để workflow LeRobot/VLA dễ chia sẻ hơn. Xem platform →

Bài toán: Tại Sao VLA Thất Bại Khi OOD?

VLA như π0 được pre-train trên hàng nghìn giờ dữ liệu manipulation. Khi deployment gặp task tương tự training, model hoạt động tốt. Nhưng khi instruction được diễn đạt khác, hoặc môi trường thay đổi nhẹ (màu sắc, vị trí object, góc camera), model có thể thất bại một cách không đoán trước được.

Vấn đề gốc rễ: Behavior cloning — cách VLA học — optimize model để copy mode trội nhất trong demonstration data. Khi gặp OOD, model bị kẹt vào action mode đó, không explore alternative behaviors có thể thành công hơn.

RL là giải pháp tự nhiên để khắc phục: RL có thể discover high-value behaviors ngoài demonstration modes. Nhưng RL fine-tuning toàn bộ VLA rất đắt — model có hàng tỷ parameter, update cần nhiều online rollout, và có nguy cơ forgetting knowledge cũ.

RL²-VLA tách biệt hai mục tiêu: giữ nguyên VLA gốc (kiến thức rộng, stable), và train một steering module nhỏ chuyên biệt để làm RL — module này chỉ cần học cách điều chỉnh action generation của VLA, không cần học lại toàn bộ perception và world model.


Kiến Trúc: Ba Module Phối Hợp

Kiến trúc RL²-VLA với 3 module SAFE, QAM và CoVer — nguồn: marmotlab/RL2-VLA
Kiến trúc RL²-VLA với 3 module SAFE, QAM và CoVer — nguồn: marmotlab/RL2-VLA

RL²-VLA gồm ba thành phần tách biệt, phối hợp với nhau tại inference time:

1. SAFE — Failure Detector

SAFE (Scalable Adaptive Failure Estimator) là một LSTM nhỏ được condition trên internal latent features của VLA action expert — không phải raw observations. Nó học cách phân biệt rollout đang đi về phía thành công hay thất bại.

Điểm độc đáo: SAFE dùng conformal prediction để calibrate threshold. Thay vì set một ngưỡng cố định, conformal prediction tự động điều chỉnh threshold sao cho với xác suất 1-α, bất kỳ rollout thành công nào cũng không bị nhầm thành failure. Điều này tránh việc kích hoạt steering không cần thiết khi VLA đang hoạt động tốt.

Quá trình calibration:

  1. Thu thập rollouts từ base VLA (60% train, 40% validation)
  2. Train SAFE để predict failure score ∈ [0, 1] cho mỗi timestep
  3. Dùng validation rollouts thành công để fit conformal bands
  4. Per-task thresholds đảm bảo false positive rate thấp

2. QAM — RL Steering Policy

QAM (Q-learning with Adjoint Matching) là core của phương pháp. Đây là một lightweight flow-matching policy được train bằng offline RL, conditioned trên latent representations từ VLA action expert.

Vấn đề kỹ thuật quan trọng: Tại sao cần "Adjoint Matching"? Flow-matching policy generate action qua multi-step ODE integration. Backpropagation xuyên qua nhiều bước này numerically unstable — gradient explode hoặc vanish. Adjoint Matching giải quyết bằng cách tính gradient thông qua reverse ODE, ổn định hơn và không cần backprop qua toàn bộ forward pass.

Composition tại inference: Khi SAFE detect failure, velocity của VLA action expert và velocity của QAM steering policy được compose:

v̂ = w · v_VLA + (1 - w) · v_RL

Trong đó weight w được sample từ Gaussian(μ=0.5, σ=0.25), đảm bảo VLA vẫn giữ vai trò chính nhưng có sự dẫn dắt từ RL policy.

Tại sao latent space? Thay vì steer ở observation space hay action space raw, QAM operate trong expressive latent của VLA. Những latent này đã encode context phong phú (scene, instruction, robot state) — RL policy có thể condition trên context đó để tạo ra guidance phù hợp hơn.

3. CoVer — Action Verifier

CoVer (Contrastive Verifier) là một external reward model học để score candidate actions. Khi cần chọn giữa nhiều actions được generate (multi-sample), CoVer rank chúng và chọn action tốt nhất để execute.

CoVer được train với contrastive learning: với cùng một state, action dẫn đến success được rank cao hơn action dẫn đến failure. Model này không cần retrain khi switch VLA backbone — nó chỉ score actions, không generate.


Test-Time Scaling Laws

Một phát hiện quan trọng của paper là scaling laws khác biệt giữa success và failure states:

  • Khi VLA đang thành công, thêm samples/diversity không giúp ích — thậm chí có thể làm hỏng action chính xác đang được generate.
  • Khi VLA đang thất bại, thêm samples/diversity cải thiện đáng kể — explore alternative modes tìm ra path thành công.

Đây là lý do tại sao adaptive steering (chỉ steer khi failure predicted) outperform always-on steering trên nhiều benchmark. SAFE detect được moment nào cần can thiệp và moment nào nên để VLA tự hoạt động.

Kết quả scaling: dùng 8 cách rephrase instruction × 5 samples/rephrase đạt +18.6% improvement — nhưng chỉ khi kết hợp với adaptive triggering.


Cài Đặt và Setup

Yêu cầu hệ thống

  • Linux Ubuntu 20.04/22.04
  • Python 3.10
  • GPU với CUDA (khuyến nghị H100, A6000, hoặc RTX 5090; có thể chạy evaluate trên RTX 3090+)
  • VRAM: ≥24GB để load π0 base model

Clone repository

git clone --recurse-submodules https://github.com/marmotlab/RL2-VLA.git
cd RL2-VLA

Lưu ý --recurse-submodules — repo dùng submodule cho SAFE component, thiếu flag này sẽ thiếu code.

Tạo conda environment

conda create -n rl2 python=3.10
conda activate rl2

Install dependencies

# Install tất cả dependencies cho SIMPLER simulation + π0 integration
bash RL2_CoVer_VLA/env_simpler_pi.sh

Script này install: lerobot, simpler-env, pytorch, diffusers, và các dependencies cần thiết. Quá trình mất 10-20 phút tuỳ tốc độ mạng.

Download pretrained models

# Download QAM steering policy từ HuggingFace
huggingface-cli download rl2-vla/qam-steering-policy --local-dir ./checkpoints/qam

# Download CoVer verifier
huggingface-cli download cover-vla/cover-vla-bridge --local-dir ./checkpoints/cover

SAFE failure detector đã được include trong submodule, không cần download riêng.


Training Pipeline (Nếu Muốn Train Từ Đầu)

RL²-VLA hỗ trợ train trên custom dataset. Pipeline gồm 4 bước:

Bước 1: Extract VLA Latents

# Extract latent representations từ frozen VLA action expert
# Chạy trên BridgeV2 hoặc DROID dataset
python RL2_CoVer_VLA/extract_latents.py \
    --dataset bridge_v2 \
    --vla_checkpoint path/to/pi0 \
    --output_dir data/latents/

Latents được extract từ internal hidden states của action expert transformer — không phải từ visual encoder hay language encoder. Đây là những features cao cấp đã capture task context.

Bước 2: Train QAM Steering Policy

# Train offline RL flow-matching policy conditioned trên latents
python RL2_CoVer_VLA/train_qam.py \
    --latent_dir data/latents/ \
    --batch_size 256 \
    --lr 3e-4 \
    --epochs 100 \
    --output_dir checkpoints/qam/

Training dùng Q-learning với Adjoint Matching:

  • Q-values được estimate từ offline data (không cần online rollout)
  • Policy update dùng reverse ODE để tính gradient ổn định
  • Điều kiện hóa trên latent để policy "biết" context hiện tại

Trên A100, training khoảng 2-4 giờ với BridgeV2.

Bước 3: Train SAFE Failure Detector

# Collect rollouts từ base VLA để train SAFE
python RL2_CoVer_VLA/collect_rollouts.py \
    --env simpler \
    --tasks bridge_tasks \
    --n_rollouts 2000 \
    --output_dir data/rollouts/

# Train SAFE LSTM
python RL2_CoVer_VLA/train_safe.py \
    --rollout_dir data/rollouts/ \
    --train_ratio 0.6 \
    --output_dir checkpoints/safe/

Bước 4: Calibrate Conformal Prediction

# Calibrate thresholds dùng validation rollouts
python RL2_CoVer_VLA/calibrate_safe.py \
    --rollout_dir data/rollouts/ \
    --safe_checkpoint checkpoints/safe/ \
    --alpha 0.1 \  # 1-alpha = 90% coverage
    --output_dir checkpoints/safe/thresholds/

alpha=0.1 nghĩa là với 90% probability, rollout thành công sẽ không trigger false-positive steering.


Inference và Evaluation

Evaluation trên SIMPLER Benchmark

conda activate rl2

# Adaptive steering (khuyến nghị) — chỉ steer khi SAFE predict failure
bash RL2_CoVer_VLA/simpler/bashes/eval_rl2_compose_adaptive.sh

# Always-on steering — steer mọi timestep (baseline so sánh)
bash RL2_CoVer_VLA/simpler/bashes/eval_rl2_compose_always.sh

# Rephrase-only baseline — không có RL steering
bash RL2_CoVer_VLA/simpler/bashes/eval_rephrase.sh

Tổng hợp kết quả

python RL2_CoVer_VLA/simpler/bashes/summarize_logs.py \
    --log_dir logs/simpler_eval/

Output là bảng success rate per task, plus overall average.

Integration vào custom robot

Để deploy trên robot thực (ví dụ: PiperX arm):

import torch
from rl2_vla import RL2Pipeline

# Load pipeline
pipeline = RL2Pipeline.from_pretrained(
    vla_checkpoint="path/to/pi0",
    qam_checkpoint="checkpoints/qam/",
    safe_checkpoint="checkpoints/safe/",
    cover_checkpoint="checkpoints/cover/",
    adaptive=True  # chỉ steer khi cần
)

# Inference loop
obs = robot.get_observation()
instruction = "pick up the red cup and place it in the bowl"

action = pipeline.predict(
    observation=obs,
    instruction=instruction,
    n_samples=5  # số action samples để CoVer rank
)

robot.execute(action)

Kết Quả Benchmark

SIMPLER Simulation

SIMPLER là simulation benchmark chạy trên physics-based environment với các tasks như pick-and-place, push, rotate. RL²-VLA được test với OOD instruction prompts — cùng task nhưng mô tả khác với training data.

Phương pháp Success Rate So với Base VLA
π0 base (no steering) baseline —
Rephrase only +5.2% —
RL²-VLA (always-on) +10.1% avg +14.7% max task
RL²-VLA (adaptive) +14.7% avg Tốt nhất

Adaptive steering nhất quán tốt hơn always-on vì tránh perturbation khi VLA đang chính xác.

PolaRiS Benchmark

PolaRiS test với π0.5 (model mạnh hơn π0) trên OOD scenarios khó hơn:

Phương pháp Success Rate Improvement
PolaRiS (π0.5 base) baseline
RL²-VLA adaptive +17.3% avg
Scaling (8 rephrases × 5 samples) +18.6%

Real Robot: PiperX Arm

Kết quả quan trọng nhất — test trên hardware thực. PiperX là 6-DOF desktop manipulator với wrist camera:

Phương pháp Success Rate
π0 base baseline
Rephrase baseline +9.2%
RL²-VLA adaptive +26.7% avg (+19.5% overall)

Real-robot improvement lớn hơn simulation vì OOD gap trong thực tế lớn hơn, và adaptive steering giải quyết trực tiếp vấn đề này.


So Sánh Với Các Phương Pháp Liên Quan

Có nhiều hướng để cải thiện VLA performance tại test time. RL²-VLA định vị như sau:

Phương pháp Có cần retrain? Latent steering? Adaptive? Real robot?
RL fine-tuning ✅ (tốn) ✗ ✗ Ít khi test
Test-time search ✗ ✗ ✗ ✗
Instruction rephrase ✗ ✗ ✗ ✗
RL²-VLA ✗ ✅ ✅ ✅

So sánh với ResVLA — cũng dùng residual approach nhưng tập trung vào intent anchoring thay vì failure detection. RL²-VLA mạnh hơn ở adaptive triggering và tích hợp offline RL.

So sánh với test-time compute scaling — RL²-VLA thực chất là một dạng test-time compute đặc biệt: thay vì chỉ tăng số samples, nó dùng RL-guided steering để improve quality từng sample.

Để hiểu nền tảng về VLA architecture, xem OpenVLA deep dive.


Phân Tích Ablation

Paper thực hiện ablation kỹ càng để justify từng component:

Latent conditioning vs raw observation: Dùng latent tốt hơn raw obs +4.2% — latent đã compress context phong phú hơn.

RL vs Behavior Cloning cho steering: RL-trained QAM tốt hơn BC-trained policy +4.5% — RL discover alternative behaviors ngoài demonstration modes.

Adaptive vs Always-on: Adaptive tốt hơn +2.3% trên OOD tasks và tránh degradation trên in-distribution tasks.

CoVer verifier: Thêm CoVer verification khi multi-sampling cải thiện thêm +1.8%.


Hạn Chế và Hướng Mở Rộng

Hạn chế hiện tại:

  • QAM train trên BridgeV2/DROID — performance có thể khác trên domains không liên quan
  • SAFE cần calibration rollouts — trên robot hoàn toàn mới cần thu thập thêm data
  • GPU requirement vẫn cao (≥24GB VRAM cho π0 base)

Hướng mở rộng thú vị:

  • Apply cho VLA khác (không chỉ π0): OpenVLA-OFT, RoboVLMs
  • Extend sang manipulation phức tạp hơn: bimanual, contact-rich
  • Online adaptation: update QAM dần khi collect real rollouts mới

Kết Luận

RL²-VLA là minh chứng rằng không cần retrain VLA để improve performance đáng kể. Bằng cách train một steering module nhỏ trong latent space và kích hoạt nó một cách thông minh (chỉ khi failure predicted), framework đạt +26.7% success rate trên robot thực — con số ấn tượng với chi phí compute thấp.

Kiến trúc ba thành phần SAFE + QAM + CoVer tách biệt rõ ràng: SAFE biết khi nào cần can thiệp, QAM biết cách can thiệp, CoVer biết chọn action nào. Sự phân tách này cho phép upgrade từng component độc lập khi VLA base model mới ra.

Với trend test-time compute scaling đang nổi trong robotics, RL²-VLA mở ra hướng mới: không phải scaling naive (nhiều samples hơn), mà intelligent latent steering với RL-guided diversity, adaptive triggering, và verified selection.


Bài Viết Liên Quan

  • ResVLA: Residual Diffusion Bridge để Anchor VLA Policy từ Intent
  • Robo-ValueRL: Deploy VLA-RL Precision Manipulation cho Robot Humanoid
  • Expo-FT Pi0.5: Online RL Fine-tune VLA trong 19 Phút
NT

Nguyễn Anh Tuấn

Robotics & AI Engineer. Building VnRobo — sharing knowledge about robot learning, VLA models, and automation.

Khám phá VnRobo

Fleet MonitoringROS 2 IntegrationAMR Solutions

Bài viết liên quan

NEWTutorial
FM-VLA: Force Memory Token cho VLA Contact-Rich Manipulation
vlaforce-sensingmanipulation
wholebody-vla

FM-VLA: Force Memory Token cho VLA Contact-Rich Manipulation

FM-VLA dùng VAE nén lịch sử lực thành Force Memory Tokens, giúp VLA vượt giới hạn Markovian — đếm contact, nhớ tiến trình, đạt 83.3% trên robot AgiBot G1.

31/7/202614 phút đọc
NT
NEWTutorial
Pelican-VLA 0.5 trên LeRobot 3.0
pelican-vlalerobotvla
wholebody-vla

Pelican-VLA 0.5 trên LeRobot 3.0

Hướng dẫn chạy Pelican-VLA 0.5, hiểu Bottleneck Token, chuẩn bị LeRobot 3.0, training, inference và đọc kết quả RoboTwin.

29/7/202616 phút đọc
NT
NEWNghiên cứu
ResVLA ICML 2026: Anchor VLA Policy từ Intent với Residual Bridge
vladiffusionmanipulation
wholebody-vla

ResVLA ICML 2026: Anchor VLA Policy từ Intent với Residual Bridge

ResVLA (ICML 2026): dùng Residual Diffusion Bridge anchor VLA từ intent, giảm nhiễu, tăng robustness cho robot manipulation.

27/7/202615 phút đọc
NT
VnRobo logo

Hạ tầng AI cho robot công nghiệp thế hệ mới.

Sản phẩm

  • Tính năng
  • Bảng giá
  • Kiến thức
  • Dịch vụ

Công ty

  • Về chúng tôi
  • Blog
  • Liên hệ

Pháp lý

  • Chính sách bảo mật
  • Điều khoản sử dụng

© 2026 VnRobo. Bảo lưu mọi quyền.

Được tạo với♥tại Việt Nam