VnRoboVnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
VnRobo logoVnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam
VnRoboVnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
  1. Home
  2. Blog
  3. TrajBooster: 6D Trajectory Transfer for Whole-Body VLA
wholebody-vlawholebody-vlavlawhole-bodygrootunitree-g1cross-embodimenttrajectory-retargetingimitation-learning

TrajBooster: 6D Trajectory Transfer for Whole-Body VLA

Transfer Agibot demonstrations to Unitree G1 through 6D trajectories, simulated retargeting, GR00T N1.5, and limited target-robot fine-tuning.

Nguyễn Anh TuấnOctober 8, 202614 min read
TrajBooster: 6D Trajectory Transfer for Whole-Body VLA

A wheeled robot already knows how to manipulate objects at several heights. Your new bipedal humanoid has only a few minutes of demonstrations. Can you reuse the older data without forcing the humanoid to copy another robot's joint angles?

TrajBooster uses the 6D trajectories of both wrists as a transfer interface. It passes those trajectories through simulation to generate actions compatible with Unitree G1, uses the resulting data to adapt a VLA, and then fine-tunes with real target-robot demonstrations. The useful engineering lesson is how to create new action labels for a different body.

This guide follows the TrajBooster v3 paper, dated March 19, 2026, and the official OpenHelix-Team/OpenTrajBooster implementation. The repository announces acceptance at ICRA 2026. The initial paper appeared in 2025; this is not an October 2026 paper announcement. Our focus is trajectory transfer and data contracts, extending the blog's existing coverage of robot teleoperation.

1. Why copying joint positions does not transfer a skill

Imagine two people with different arm lengths reaching for the same cup. They share a hand destination, but their shoulder and elbow angles differ. Robots add further differences: joint counts, limits, coordinate conventions, cameras, and balance requirements.

An embodiment includes those physical properties and control interfaces. Cross-embodiment learning transfers useful behavior across them. A model that treats actions as anonymous lists of numbers can learn the wrong mapping even while its training loss falls.

A 6D pose describes three degrees of freedom in position and three in orientation. That does not require exactly six stored numbers: an orientation may use a four-value quaternion or a nine-value rotation matrix. Separate the physical meaning from the serialization format.

TrajBooster transfers the paths of two end-effectors rather than the source robot's complete joint configuration. Each robot can attempt to move its hands along a shared task trajectory while using its own body. The time dimension matters as well: approach, close the hand, lift, and relocate. A single target pose does not represent that sequence.

For a beginner, the central question is whether a demonstration describes what the hands should accomplish or how one particular body accomplished it. The former can be a better transfer interface, although it still needs workspace mapping and a controller that makes the motion feasible.

If you plan to train the policy, prioritize the dataset, GPU environment, and action mapping before expanding the hardware setup.

Tool recommendations

VLA train/deploy stack

Train on cloud/workstation, then deploy optimized models to Jetson or the robot computer.

Cloud GPU for VLA / policy training Use for imitation learning, diffusion policies, RL, and robotics model fine-tuning. View cloud GPU → NVIDIA Jetson Orin NX / Orin Nano Edge deployment hardware for perception, logging, and optimized inference. View Jetson → Hugging Face / robotics dataset hosting Host datasets, checkpoints, and model cards for cleaner LeRobot/VLA workflows. View platform →

2. The real-to-sim-to-real pipeline

The pipeline contains three transitions, each addressing a different mismatch:

text
Agibot-World: cameras + language + wrist trajectories
                           |
                  Source workspace mapping
                           |
          Unitree G1 retargeting in Isaac Gym
                           |
      Source images + source language + target actions
                           |
          GR00T N1.5 post-pre-training (PPT)
                           |
       Fine-tuning with real Unitree G1 demonstrations
                           |
          VLA + worker policy -> whole-body execution

TrajBooster real-to-sim wrist tracking and source versus simulated robot views — source: TrajBooster paper, Figure 5
TrajBooster real-to-sim wrist tracking and source versus simulated robot views — source: TrajBooster paper, Figure 5

The first transition handles kinematic differences. The second builds data for adapting the model to a new action space. The last adapts the policy to the target robot's real images and states. Calling the intermediate dataset real G1 demonstrations would obscure how it was produced.

The authors call its samples heterogeneous triplets: source vision, source language, and target action. They replace the action labels while retaining source images. This is practical, but visually inconsistent: the policy can see the source robot's hands while learning to predict G1 actions.

Target-robot fine-tuning helps address that mismatch. It does not establish that camera geometry, appearance, and contact differences have disappeared. When interpreting the results, remember that the intermediate data is a bridge rather than a perfectly matched target dataset.

3. Mapping trajectories into G1's workspace

Section III-A reports an approximately 1.8-meter arm span for Agibot and a 1.2-meter span for Unitree G1. A source trajectory can therefore request a target that G1 cannot reach.

The paper aligns the x-axis through z-score statistics referenced to G1 data, scales y by 0.6667 according to arm-length ratio, and clips z to 0.15–1.25 meters. These are study-specific preprocessing choices, not universal operating limits for humanoids.

Think of this as resizing a drawing before handing it to a different machine. Converting centimeters to meters only fixes units. It does not make an oversized path reachable. Conversely, shrinking positions does not fix a wrongly oriented coordinate frame.

A useful preparation exercise is to plot one short episode before processing thousands. Label the axes, distinguish world and base frames, verify meters and radians, and compare a small motion with a result you can predict. Check orientation conventions as carefully as position. These are practical recommendations from this guide, not additional experiments reported by the authors.

Retargeting errors can otherwise become systematic training labels. A clean dataset file is not necessarily a physically correct dataset: a wrong sign or swapped hand may remain perfectly valid numerically.

4. Retargeting architecture: arm, manager, and worker

Whole-body retargeting requires more than solving inverse kinematics for two arms. Low targets may require squatting, and reaching changes the upper body's motion and disturbances experienced by the legs.

Component Responsibility Main output
Arm policy Closed-loop inverse kinematics with Pinocchio Arm joint position targets
Manager policy Choose base motion from wrist goals vx, vy, yaw rate, torso height
Worker policy Execute base commands under changing upper-body motion Targets for 12 leg joints

The worker follows a Homie-style RL approach with a curriculum that increases disturbances from upper-body motion. The retargeting README also describes locomotion training in semi-crouched postures. This matters because manipulating a low object is different from walking with stationary arms.

The manager uses heuristic-enhanced harmonized online DAgger. DAgger addresses a familiar imitation-learning problem: a learned policy visits states different from the expert's demonstrations. A policy trained only on expert states may not know how to recover after drifting.

Here, manager labels come from heuristics and privileged simulation information. Seed trajectories are augmented with height changes through PCHIP interpolation. The manager learns from its rollouts, with new demonstrations added to an aggregate dataset once every ten iterations. This balances retaining earlier experience against storage and processing costs.

Those simulator-derived labels should not be confused with sensors available to the deployed VLA. The simulation can expose information used to construct training targets that will not exist at inference time. The retargeting README documents the worker and manager as separate modules.

5. What the deployed VLA actually predicts

The VLA is based on GR00T N1.5, with an OpenWBC data configuration. It consumes images, a language instruction, and proprioception: the robot's joint state. Its flow-matching action head generates continuous action chunks.

For an intuitive picture of flow matching, imagine learning how to transform a noisy action sequence into an appropriate one, conditioned on the scene and instruction. The model does not need to produce a prose plan before generating motor commands.

The paper uses 16-timestep action chunks at 20 Hz with four denoising steps. A complete chunk spans 0.8 seconds if executed sequentially. The horizon alone does not tell you whether the deployment client executes every predicted action before collecting a fresh observation. That depends on the implementation of its control loop.

The VLA outputs arm and hand joint targets plus four base commands. A worker policy converts those base commands into leg joint targets. Whole-body behavior therefore comes from a hierarchical system. The VLA is not directly predicting torques for every actuator.

This separation also explains a debugging rule: if the hands make reasonable predictions but the body behaves poorly, inspect the base-command interface and worker before assuming the vision backbone is the problem.

6. Installation: start offline and separate environments

The commands below are adapted from the repository documentation. They have not been executed as a training run or validated on hardware for this article. Replace dataset and checkpoint paths with your own resources.

The retargeting stack recommends Ubuntu 20.04/22.04, Isaac Gym Preview 4, and Python 3.8. The bundled GR00T documentation recommends Python 3.10 and CUDA 12.4 for the VLA environment. Keep those environments separate rather than assuming one installation will support everything.

Download the archive from the official GitHub repository, extract it, and name the directory OpenTrajBooster. Then prepare the VLA environment:

bash
cd OpenTrajBooster/VLA_model/gr00t_modified_for_OpenWBC
conda create -n trajbooster-vla python=3.10
conda activate trajbooster-vla
pip install --upgrade setuptools
pip install -e '.[base]'
pip install --no-build-isolation flash-attn==2.7.1.post4
python scripts/gr00t_finetune.py --help

Check the driver, PyTorch build, and CUDA environment before investigating FlashAttention compilation failures. If imports fail while displaying command help, solve that first instead of downloading a large dataset. The README documents a specific OpenCV CV_8U problem; apply its workaround when you encounter that error rather than reinstalling packages indiscriminately.

Starting from the released PPT checkpoint avoids repeating the entire intermediate training stage. The project also releases the Agibot-to-G1 retargeted dataset. Its README estimates roughly 30 GB for the dataset and 6 GB for the checkpoint. Actual downloads depend on the revision and files selected.

7. The data contract: 40 state dimensions, 32 actions

The converter's modality.json is essential reading before training:

Group Raw state dimensions Action dimensions
Left and right arms 14 14
Left and right hands 14 14
Left and right legs 12 0
Base motion 0 4
Total 40 32

The model observes leg states without directly predicting twelve leg action values. The final four action dimensions are worker commands, not four extra joint angles. Raw states can also undergo feature transforms before entering the network; the stored dataset dimension and the processed model feature dimension are different concepts.

The openwbc_g1 configuration expects three video streams: ego_view, wrist_left, and wrist_right. It uses the current observation and sixteen future actions. Confirm camera order, annotation names, and dimension ranges instead of relying on a generic LeRobot dataset loader to infer them.

The top-level README provides this three-view conversion command:

bash
cd OpenTrajBooster/OpenWBC_to_Lerobot
python convert_3views_to_lerobot.py \
  --input_dir /data/g1_raw \
  --output_dir /data/g1_lerobot \
  --dataset_name pick_toy \
  --robot_type g1 \
  --fps 30

The converter documentation also includes a single-view example using convert_to_lerobot.py. Do not silently replace the three-view script with that command. Likewise, the example conversion rate of 30 FPS and inference rate of 20 Hz require an explicit timing policy. They do not automatically represent the same action interval.

Before a long training run, inspect consecutive frames and actions from several episodes. Check whether the image timestamp matches the command, left and right hands are correctly assigned, and the language describes the recorded task. Split training and validation by episode so adjacent frames from one demonstration do not appear in both sets.

8. Training: PPT followed by target-robot fine-tuning

The study uses 1,960 episodes across 176 tasks, producing roughly 35 hours of retargeted data. PPT runs for 60,000 steps on two A100 80 GB GPUs with batch size 128. Target fine-tuning uses approximately ten minutes of data, comprising 28 episodes, with one A100, batch size sixteen, and 3,000 steps.

The crucial interpretation is that ten minutes refers to additional target-domain demonstrations. Pretrained weights, source demonstrations, retargeting, and GPU training still exist upstream. It is not the total data or compute budget of the project.

An example target fine-tuning command, adapted from the VLA README, starts from a locally downloaded PPT model:

bash
cd OpenTrajBooster/VLA_model/gr00t_modified_for_OpenWBC
python scripts/gr00t_finetune.py \
  --base-model-path /models/PPTmodel4UnitreeG1 \
  --dataset-path /data/g1_lerobot \
  --output-dir ./save/pick_toy_ppt \
  --data-config openwbc_g1 \
  --embodiment-tag new_embodiment \
  --num-gpus 1 \
  --batch-size 16 \
  --max_steps 3000 \
  --report-to tensorboard

Matching the number of steps does not reproduce the paper by itself. Dataset splits, camera setup, normalization, frozen modules, and code revision must also match. Begin with a short run that confirms all modalities load and checkpoints save before committing to the full schedule.

The two-stage procedure is worth testing against a controlled baseline. Keep your target demonstrations fixed, train one model with the PPT initialization and another without it, and evaluate both under the same conditions. Otherwise, additional data collection or a changed task can obscure whether transfer actually helped.

9. Results: separate complete and partial success

Four TrajBooster evaluation tasks at different manipulation heights — source: TrajBooster paper, Figure 6
Four TrajBooster evaluation tasks at different manipulation heights — source: TrajBooster paper, Figure 6

Table II evaluates each task over ten trials. The table below reports complete task success separately from partial outcomes:

Task PPT + 3K steps No PPT + 3K steps No PPT + 10K steps
Pick Mickey Mouse 100% 0% 80%
Store Toys 70% 0% 30%
Clean the Table 70% 0% 70%
Pick Orange & Place 10% 0% 0%

The orange task exposes an important limitation. The PPT model has an additional 50% of trials with successful grasping but failed placement. Adding those to report 60% task success would misrepresent the result. With only ten trials, one outcome changes the reported rate by ten percentage points.

Table III moves the toy to positions absent from the teleoperation demonstrations. The PPT model reaches 80% success, while the no-PPT model trained for 10K steps reaches 0%. The authors' trajectory analysis suggests the baseline tends to repeat the demonstrated path, whereas the PPT model adapts its approach.

Training loss curves comparing action-space adaptation approaches — source: TrajBooster project page
Training loss curves comparing action-space adaptation approaches — source: TrajBooster project page

Loss curves are useful for tracking optimization, but they do not replace task evaluation. A model that closely predicts the demonstrated motion may fail when the object moves sideways. Include both task success and changed object positions in evaluation, and record where each attempt failed.

The “Pass the Water” demonstration is zero-shot with respect to target-robot teleoperation for that task. The task appears in PPT source data but not in the G1 fine-tuning set. It is not a skill absent from every stage of training.

10. Inference and deployment checks

Start with offline evaluation on separate episodes before connecting actuators:

bash
python scripts/eval_policy.py \
  --plot \
  --model_path ./save/pick_toy_ppt/checkpoint-3000 \
  --dataset-path /data/g1_validation \
  --embodiment-tag new_embodiment \
  --data-config openwbc_g1 \
  --modality-keys base_motion left_hand right_hand

Look for abrupt base-command changes, unreasonable requested heights, switched hand labels, and isolated spikes. An average error can hide a short interval with poor behavior. Validation episodes should use the same schema as training while remaining independent demonstrations.

Real deployment additionally requires the image server, HomieDeploy, Unitree SDK2, and the correct network setup. After validating the low-level stack separately, the repository's VLA interface takes this form:

bash
python scripts/G1_inference.py \
  --arm=G1_29 \
  --hand=dex3 \
  --model-path ./save/pick_toy_ppt/checkpoint-3000 \
  --goal pick_toy \
  --frequency 20 \
  --vis \
  --filt

The README recommends an initial pose close to the teleoperation data and 20 Hz operation. Inspect outputs offline, verify the controller in simulation, and then conduct supervised hardware trials with an appropriate stopping mechanism. A filtering flag does not establish that every controller interaction is correct.

For a first experiment, use one object and two manipulation heights. Compare the same target dataset with and without PPT, then move the object to a new position. Record perception, grasping, placement, and body-stability failures separately. That breakdown tells you whether the next improvement belongs in the vision input, demonstration coverage, or worker controller.

TrajBooster offers a concrete recipe for turning another robot's data into a useful foundation for whole-body VLA. Its paper still identifies limitations in hand precision, visual consistency, available loco-manipulation data, and experiment scale. A narrow, measurable transfer experiment is a sensible starting point before attempting a household generalist.

Related Posts

  • OpenWBC: VR teleoperation and G1 data collection
  • Ψ₀: Whole-body VLA architecture
  • OpenHLM: Whole-body humanoid loco-manipulation
NT

Nguyễn Anh Tuấn

Robotics & AI Engineer. Building VnRobo — sharing knowledge about robot learning, VLA models, and automation.

Explore VnRobo

Fleet MonitoringROS 2 IntegrationAMR Solutions
← Previous
VLA-ACL: Train Visual Token Pruning on LIBERO
Next →
UFO and TeCH: Humanoid WBC with Unsupervised RL

Related Posts

Tutorial
FluxVLA Engine: Hands-On Guide to the One-Stop VLA Platform
vlalerobotpi0
wholebody-vla

FluxVLA Engine: Hands-On Guide to the One-Stop VLA Platform

LimX Dynamics' open-source FluxVLA Engine covers the full pipeline from data to real-robot deployment, supporting Pi0, GR00T, and OpenVLA in one unified framework.

9/17/20269 min read
NT
Tutorial
DSPv2: Dense Policy for Whole-Body Mobile Manipulation
vlawbcwhole-body
wholebody-vla

DSPv2: Dense Policy for Whole-Body Mobile Manipulation

DSPv2 fuses 3D point clouds with multi-view DINOv2 semantics for generalizable whole-body mobile manipulation — complete guide to setup, training, and inference.

9/15/202614 min read
NT
Research
IntentVLA: Fixing Observation Aliasing in Robot Manipulation
vlamanipulationimitation-learning
wholebody-vla

IntentVLA: Fixing Observation Aliasing in Robot Manipulation

IntentVLA encodes visual history into a compact intent token to resolve observation aliasing — the root cause of unstable action chunks in VLA fine-tuning.

9/13/202613 min read
NT
VnRobo logoVnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam