Hầu hết các mô hình VLA (Vision-Language-Action) hiện nay dạy robot chuyển động đúng hướng, đến đúng điểm — nhưng hoàn toàn bỏ qua một câu hỏi thiết yếu: dùng bao nhiêu lực? Khi robot humanoid phải nhấc hộp nặng, đẩy đồ vật dọc kệ, hay lau bàn — mỗi task đòi hỏi mức lực tiếp xúc khác nhau, và không có thông tin lực, controller sẽ hoặc là quá nhẹ tay (trượt, thất bại) hoặc quá mạnh tay (hỏng vật, ngã). Đây chính xác là khoảng trống mà Opt2VLA của nhóm Georgia Tech lấp đầy.
Opt2VLA: Force-Aware Vision-Language-Action for Contact-Rich Humanoid Whole-Body Manipulation — Fukang Liu*, Yipu Chen*, Jaehwi Jang, Danfei Xu, Zsolt Kira, Ye Zhao — Georgia Institute of Technology, arXiv 2609.23968 (tháng 9/2026).
"Opt2" có nghĩa gì?
Tên "Opt2" trong Opt2VLA đến từ Optimization-to — cụ thể là Whole-Body Trajectory Optimization (TO). Đây là cách đặt tên nhất quán với paper trước của nhóm là Opt2Skill (RA-L 2025). Điểm khác biệt cốt lõi: thay vì dùng dữ liệu teleoperation con người (vốn không ghi nhận lực), Opt2VLA dùng solver DDP (Differential Dynamic Programming) trong Crocoddyl để tạo ra trajectory tối ưu với lực tiếp xúc được tính toán chính xác từ vật lý. Đây là "nguồn sự thật" về lực mà không bộ dữ liệu human demo nào có thể cung cấp.
Vấn đề cốt lõi: VLA thiếu chiều lực
Hãy tưởng tượng bạn dạy một học sinh cách lau bàn bằng cách chỉ mô tả quỹ đạo tay di chuyển — không bao giờ nói "ấn mạnh hơn" hay "nhẹ thôi." Kết quả sẽ là học sinh đó lau bàn với lực ngẫu nhiên, không đoán được. Đây đúng là tình trạng của VLA hiện tại với manipulation task.
Cụ thể, các VLA state-of-the-art như π0, GR00T-N1 đưa ra:
- Output: 3D waypoints cho end-effector (tay, chân)
- Thiếu: Reference lực tiếp xúc
Downstream whole-body controller (WBC) nhận waypoint đó và phải tự đoán cần bao nhiêu lực — không có căn cứ gì. Với free-space manipulation, điều này chấp nhận được. Nhưng với contact-rich tasks (push, wipe, press), đây là điểm mù chết người.
Opt2VLA giải quyết bằng cách mở rộng action space của VLA để bao gồm cả force reference, và dạy VLA học cách ánh xạ từ ngôn ngữ ("nhẹ nhàng," "vừa phải," "mạnh") sang mức lực cụ thể.
Kiến trúc hai tầng
Opt2VLA là hệ thống phân cấp gồm 2 tầng hoạt động nối tiếp nhau:

Tầng 1: VLA Policy (GR00T-N1.7, fine-tuned)
Model nền tảng là GR00T-N1.7 của NVIDIA, được fine-tune với input và output mở rộng.
Input mở rộng:
- Language instruction ("Pick up the box gently")
- Egocentric RGB camera
- Proprioceptive state: joint position, joint velocity, base pose, base velocity
- MỚI: Joint torque history (10 timestep gần nhất, cách 200ms) — đây là tín hiệu cảm biến lực gián tiếp, không cần F/T sensor đắt tiền
Output mở rộng:
- 3D waypoints cho tay trái, tay phải, chân (standard WBC interface)
- MỚI: Continuous contact force reference cho điểm tiếp xúc chính
Tầng 2: Force-Conditioned RL Controller (FCT)
RL policy với kiến trúc asymmetric actor-critic. Nhận waypoints + force reference từ VLA, thực thi thành joint torques. Paper so sánh 3 biến thể:
| Biến thể | Mô tả | Force tracking? |
|---|---|---|
| MO (Motion-Only) | Baseline, không có force | ❌ |
| FC (Force-Conditioned) | Nhận force reference, có reward lực | ✅ |
| FCT (Force-Conditioned + Torque supervision) | Phương pháp đầy đủ, thêm TO-derived torque supervision khi training | ✅✅ |
FCT là phương pháp được đề xuất. Torque supervision chỉ dùng khi training (privileged learning), không cần khi inference.
Pipeline training 4 bước
Đây là phần kỹ thuật sâu nhất của paper.
Bước 1: Sinh dữ liệu bằng Trajectory Optimization
Solver: Crocoddyl với DDP (Differential Dynamic Programming).
Bài toán tối ưu bao gồm:
- Floating-base full-body dynamics của Digit humanoid (48 kg, 30 DOF)
- Contact Jacobians và contact wrench forces
- Coulomb friction cone constraints
- Joint position/velocity/torque limits
Với mỗi task configuration (vị trí vật, hướng vật, mức lực), DDP tạo ra trajectory tối ưu với nhãn lực chính xác. Đây là điều mà motion capture hay teleoperation không thể cung cấp — dữ liệu lực được tính từ vật lý, không phải đo từ sensor con người.
Bước 2: Train RL controllers (FCT)
Reward function của FCT controller (từ Table IV paper):
| Thành phần | Trọng số |
|---|---|
| Joint position tracking | 3 |
| Base position, orientation | 3 mỗi cái |
| Base linear/angular velocity | 3 mỗi cái |
| End-effector position | 3 |
| Contact force tracking | 2 |
| Joint torque tracking (TO-derived) | 2 |
| Action rate smoothness penalty | −3 |
| Torque magnitude penalty | −0.3 |
| Joint acceleration penalty | −10⁻⁵ |
Torque tracking reward là "privileged supervision" — nó dạy controller dùng joint torque gần với trajectory tối ưu, nhưng signal này bị bỏ đi hoàn toàn khi deploy (inference không cần TO data).
Bước 3: Thu thập dataset simulation
Rollout các FCT controller đã train trong simulation với cấu hình task đa dạng. Mỗi sample bao gồm:
- Visual observations (egocentric RGB)
- Language annotation mô tả mức lực ("gently lift," "firmly push," "strongly wipe")
- Proprioceptive state + torque history
- Ground truth motion goal + force reference (label)
Bước 4: Fine-tune GR00T-N1.7
GR00T-N1.7 được fine-tune trên dataset simulation trên, học cách từ (vision, language, proprioception) → (motion waypoints + force reference). Dataset này sẽ được nhóm tác giả release công khai.
Demo hardware trên robot Digit
Opt2VLA được validate trên Digit humanoid của Agility Robotics:
- Khối lượng: ~48 kg
- DOF: 30 tổng, 20 actuated joints
- Đây là nền tảng hardware duy nhất được test trong paper
Ba tasks contact-rich được demo:
Phân loại mức lực theo ngôn ngữ
| Mức lực | Ngôn ngữ | Dải lực (N) | Task |
|---|---|---|---|
| Gentle | "gently," "softly" | 0–9 N | Box pickup |
| Firm | "firmly," "steadily" | 5–16 N | Shelf-box push |
| Strong | "strongly," "firmly press" | 12–22 N | Surface wiping |
Robot chuyển đổi mức lực theo thời gian thực khi nhận lệnh ngôn ngữ mới — đây là demo closed-loop language-conditioned force modulation đầu tiên trên humanoid.
Kết quả thực nghiệm
Simulation (90 episodes, 3 tasks × 3 mức lực)
| Metric | Giá trị |
|---|---|
| Overall task success rate | 82.2% |
| Force MAE — box pickup (FCT) | 0.9 N |
| Force MAE — shelf push (FCT) | 4.2 N |
| Force MAE — surface wipe (FCT) | 1.7 N |
Ablation trên task surface wiping:
- MO (baseline): không phân biệt được mức lực, mọi level nhìn giống nhau
- FC: 2.4 ± 2.9 N force error
- FCT (proposed): 1.7 ± 2.1 N force error — cải thiện ~29%
Hardware (Digit, 45 trials)

| Task | Mean Absolute Force Error |
|---|---|
| Box pickup | 1.7 N |
| Shelf-box push | 3.9 N |
| Surface wiping | 7.0 N |
Surface wiping có sai số lớn nhất vì đây là task yêu cầu phân phối lực phức tạp nhất (cloth deformation, surface irregularity). Dù vậy, 7.0 N trong dải 12–22 N vẫn cho kết quả manipulation có kiểm soát, so với MO baseline không thể phân biệt mức lực.
Cách theo dõi và reproduce
Tại thời điểm viết bài (tháng 9/2026), nhóm tác giả chưa release code và chưa release dataset (paper ghi "will publicly release"). Tuy nhiên, bạn có thể:
Setup môi trường
# Isaac Lab cho simulation environment
git clone https://github.com/isaac-sim/IsaacLab.git
cd IsaacLab
./isaaclab.sh --install
# Crocoddyl cho trajectory optimization
pip install crocoddyl
# GR00T-N1.7 (NVIDIA)
pip install gr00t
# Hoặc từ HuggingFace: nvidia/GR00T-N1.7-3B
Reproduce pipeline conceptually
# Bước 1: TO data generation (Crocoddyl)
import crocoddyl
import numpy as np
# Setup OCP (Optimal Control Problem) cho contact task
state = crocoddyl.StateMultibody(robot_model)
actuation = crocoddyl.ActuationModelFloatingBase(state)
# Contact model (ví dụ wiping task)
contact_model = crocoddyl.ContactModelMultiple(state, actuation.nu)
contact_6d = crocoddyl.ContactModel6D(
state,
frame_id,
pinocchio.SE3.Identity(),
actuation.nu,
np.array([0., 50.]) # baumgarte gains
)
contact_model.addContact("contact_ee", contact_6d)
# Cost: lực tiếp xúc theo ngữ cảnh task
force_ref = np.array([0., 0., 15.0, 0., 0., 0.]) # "strong" = 15N
force_residual = crocoddyl.ResidualModelContactForce(
state, frame_id, pinocchio.Force(force_ref), 6, actuation.nu
)
force_cost = crocoddyl.CostModelResidual(state, force_residual)
# Solve với DDP
problem = crocoddyl.ShootingProblem(x0, running_models, terminal_model)
ddp = crocoddyl.SolverDDP(problem)
ddp.solve([], [], 300)
# Lấy optimal force trajectory
optimal_forces = [m.differential.contacts.contacts["contact_ee"].f
for m in ddp.problem.runningModels]
# Bước 2: FCT RL training (Isaac Lab)
# Reward function với force tracking
def compute_rewards(env):
# Standard motion tracking
pos_reward = torch.exp(-joint_pos_error.pow(2).sum(-1) / 0.25)
# Contact force tracking — core contribution
force_error = (measured_contact_force - reference_force).norm(dim=-1)
force_reward = torch.exp(-force_error.pow(2) / 4.0)
# Torque supervision từ TO (privileged, chỉ training)
torque_error = (joint_torques - to_reference_torques).pow(2).sum(-1)
torque_reward = torch.exp(-torque_error / 10.0)
return {
"pos": 3.0 * pos_reward,
"force": 2.0 * force_reward,
"torque_supervised": 2.0 * torque_reward, # drop khi inference
"smoothness": -3.0 * action_rate_penalty,
}
# Bước 4: Fine-tune GR00T-N1.7
from gr00t.model import GR00TN17
from gr00t.data import ForceAwareDataset
# Dataset gồm (vision, language, proprioception, torque_hist) → (waypoints, force_ref)
dataset = ForceAwareDataset(
data_dir="opt2vla_sim_dataset/",
include_torque_history=True,
torque_history_steps=10,
torque_history_interval_ms=200,
)
model = GR00TN17.from_pretrained("nvidia/GR00T-N1.7-3B")
# Mở rộng action head để predict force
model.extend_action_space(force_dim=6) # 6-DOF wrench
trainer.fine_tune(
model=model,
dataset=dataset,
epochs=50,
lr=1e-4,
)
Inference
# Inference pipeline (2 tầng)
import torch
def opt2vla_inference(rgb_obs, language_cmd, joint_state, torque_history):
"""
Args:
rgb_obs: (H, W, 3) egocentric camera
language_cmd: str, e.g. "Wipe the table firmly"
joint_state: (N_joints,) positions + velocities + base pose/vel
torque_history: (10, N_joints) joint torques, 200ms intervals
Returns:
waypoints: (T, 12) end-effector waypoints (both hands + feet)
force_ref: (6,) contact wrench reference
"""
# Tầng 1: VLA predict motion + force
with torch.no_grad():
waypoints, force_ref = vla_model(
rgb_obs, language_cmd, joint_state, torque_history
)
# Tầng 2: FCT controller execute với force reference
joint_actions = fct_controller(
current_state=joint_state,
target_waypoints=waypoints,
target_force=force_ref, # force reference từ VLA
# NOTE: không cần torque supervision khi inference
)
return joint_actions
So sánh với các approach liên quan
| Paper | Force source | Force type | Robot |
|---|---|---|---|
| Opt2VLA | Trajectory optimization | Language-conditioned, continuous | Digit humanoid |
| FM-VLA | F/T sensor history | Memory-based, reactive | Tabletop arm |
| CARE | Failure detection | Binary (success/fail) | Simulated arm |
| Force-VLA-RL | Sim RL rollout | Value-calibrated | Tabletop arm |
Điểm độc đáo của Opt2VLA: kết hợp vật lý (TO) với ngôn ngữ — người dùng chỉ cần nói "nhẹ" hay "mạnh", không cần specify con số. Đây là giao diện người dùng tự nhiên nhất trong các approach trên.
Hạn chế và hướng phát triển
Hạn chế hiện tại:
- Chỉ test 1 robot (Digit) — chưa biết performance trên Unitree G1, Boston Dynamics Atlas, hay robot arm thông thường
- Dataset chưa release — không thể reproduce đầy đủ
- 3 tasks khá đơn giản — chưa test với dexterous manipulation phức tạp (screw, pouring, cutting)
- Single contact point — force reference chỉ cho 1 điểm tiếp xúc chính, chưa xử lý multi-contact
Hướng phát triển:
- Mở rộng sang robot arm thông thường (UR5, Franka)
- Multi-contact force prediction
- Online adaptation khi vật thể có properties khác dự đoán
- Kết hợp với tactile sensor để đóng vòng lặp lực thật sự
Tại sao đây là bước tiến quan trọng
VLA cho humanoid đang ở giai đoạn mà kiểm soát chuyển động (motion control) đã khá tốt — nhưng kiểm soát lực (force control) vẫn là khoảng trống lớn. Opt2VLA là paper đầu tiên:
- Đưa force reference vào action space của VLA cho humanoid whole-body manipulation
- Dùng trajectory optimization để tạo ground-truth force labels — không cần F/T sensor đắt tiền
- Conditioned bằng ngôn ngữ — người dùng điều chỉnh lực qua từ ngữ tự nhiên, không cần specify số
Với dataset sắp được release và GR00T-N1.7 làm backbone, Opt2VLA có thể trở thành baseline standard cho force-aware VLA research trong những năm tới.



