A wheeled robot already knows how to manipulate objects at several heights. Your new bipedal humanoid has only a few minutes of demonstrations. Can you reuse the older data without forcing the humanoid to copy another robot's joint angles?
TrajBooster uses the 6D trajectories of both wrists as a transfer interface. It passes those trajectories through simulation to generate actions compatible with Unitree G1, uses the resulting data to adapt a VLA, and then fine-tunes with real target-robot demonstrations. The useful engineering lesson is how to create new action labels for a different body.
This guide follows the TrajBooster v3 paper, dated March 19, 2026, and the official OpenHelix-Team/OpenTrajBooster implementation. The repository announces acceptance at ICRA 2026. The initial paper appeared in 2025; this is not an October 2026 paper announcement. Our focus is trajectory transfer and data contracts, extending the blog's existing coverage of robot teleoperation.
1. Why copying joint positions does not transfer a skill
Imagine two people with different arm lengths reaching for the same cup. They share a hand destination, but their shoulder and elbow angles differ. Robots add further differences: joint counts, limits, coordinate conventions, cameras, and balance requirements.
An embodiment includes those physical properties and control interfaces. Cross-embodiment learning transfers useful behavior across them. A model that treats actions as anonymous lists of numbers can learn the wrong mapping even while its training loss falls.
A 6D pose describes three degrees of freedom in position and three in orientation. That does not require exactly six stored numbers: an orientation may use a four-value quaternion or a nine-value rotation matrix. Separate the physical meaning from the serialization format.
TrajBooster transfers the paths of two end-effectors rather than the source robot's complete joint configuration. Each robot can attempt to move its hands along a shared task trajectory while using its own body. The time dimension matters as well: approach, close the hand, lift, and relocate. A single target pose does not represent that sequence.
For a beginner, the central question is whether a demonstration describes what the hands should accomplish or how one particular body accomplished it. The former can be a better transfer interface, although it still needs workspace mapping and a controller that makes the motion feasible.
If you plan to train the policy, prioritize the dataset, GPU environment, and action mapping before expanding the hardware setup.
2. The real-to-sim-to-real pipeline
The pipeline contains three transitions, each addressing a different mismatch:
Agibot-World: cameras + language + wrist trajectories
|
Source workspace mapping
|
Unitree G1 retargeting in Isaac Gym
|
Source images + source language + target actions
|
GR00T N1.5 post-pre-training (PPT)
|
Fine-tuning with real Unitree G1 demonstrations
|
VLA + worker policy -> whole-body execution

The first transition handles kinematic differences. The second builds data for adapting the model to a new action space. The last adapts the policy to the target robot's real images and states. Calling the intermediate dataset real G1 demonstrations would obscure how it was produced.
The authors call its samples heterogeneous triplets: source vision, source language, and target action. They replace the action labels while retaining source images. This is practical, but visually inconsistent: the policy can see the source robot's hands while learning to predict G1 actions.
Target-robot fine-tuning helps address that mismatch. It does not establish that camera geometry, appearance, and contact differences have disappeared. When interpreting the results, remember that the intermediate data is a bridge rather than a perfectly matched target dataset.
3. Mapping trajectories into G1's workspace
Section III-A reports an approximately 1.8-meter arm span for Agibot and a 1.2-meter span for Unitree G1. A source trajectory can therefore request a target that G1 cannot reach.
The paper aligns the x-axis through z-score statistics referenced to G1 data, scales y by 0.6667 according to arm-length ratio, and clips z to 0.15–1.25 meters. These are study-specific preprocessing choices, not universal operating limits for humanoids.
Think of this as resizing a drawing before handing it to a different machine. Converting centimeters to meters only fixes units. It does not make an oversized path reachable. Conversely, shrinking positions does not fix a wrongly oriented coordinate frame.
A useful preparation exercise is to plot one short episode before processing thousands. Label the axes, distinguish world and base frames, verify meters and radians, and compare a small motion with a result you can predict. Check orientation conventions as carefully as position. These are practical recommendations from this guide, not additional experiments reported by the authors.
Retargeting errors can otherwise become systematic training labels. A clean dataset file is not necessarily a physically correct dataset: a wrong sign or swapped hand may remain perfectly valid numerically.
4. Retargeting architecture: arm, manager, and worker
Whole-body retargeting requires more than solving inverse kinematics for two arms. Low targets may require squatting, and reaching changes the upper body's motion and disturbances experienced by the legs.
| Component | Responsibility | Main output |
|---|---|---|
| Arm policy | Closed-loop inverse kinematics with Pinocchio | Arm joint position targets |
| Manager policy | Choose base motion from wrist goals | vx, vy, yaw rate, torso height |
| Worker policy | Execute base commands under changing upper-body motion | Targets for 12 leg joints |
The worker follows a Homie-style RL approach with a curriculum that increases disturbances from upper-body motion. The retargeting README also describes locomotion training in semi-crouched postures. This matters because manipulating a low object is different from walking with stationary arms.
The manager uses heuristic-enhanced harmonized online DAgger. DAgger addresses a familiar imitation-learning problem: a learned policy visits states different from the expert's demonstrations. A policy trained only on expert states may not know how to recover after drifting.
Here, manager labels come from heuristics and privileged simulation information. Seed trajectories are augmented with height changes through PCHIP interpolation. The manager learns from its rollouts, with new demonstrations added to an aggregate dataset once every ten iterations. This balances retaining earlier experience against storage and processing costs.
Those simulator-derived labels should not be confused with sensors available to the deployed VLA. The simulation can expose information used to construct training targets that will not exist at inference time. The retargeting README documents the worker and manager as separate modules.
5. What the deployed VLA actually predicts
The VLA is based on GR00T N1.5, with an OpenWBC data configuration. It consumes images, a language instruction, and proprioception: the robot's joint state. Its flow-matching action head generates continuous action chunks.
For an intuitive picture of flow matching, imagine learning how to transform a noisy action sequence into an appropriate one, conditioned on the scene and instruction. The model does not need to produce a prose plan before generating motor commands.
The paper uses 16-timestep action chunks at 20 Hz with four denoising steps. A complete chunk spans 0.8 seconds if executed sequentially. The horizon alone does not tell you whether the deployment client executes every predicted action before collecting a fresh observation. That depends on the implementation of its control loop.
The VLA outputs arm and hand joint targets plus four base commands. A worker policy converts those base commands into leg joint targets. Whole-body behavior therefore comes from a hierarchical system. The VLA is not directly predicting torques for every actuator.
This separation also explains a debugging rule: if the hands make reasonable predictions but the body behaves poorly, inspect the base-command interface and worker before assuming the vision backbone is the problem.
6. Installation: start offline and separate environments
The commands below are adapted from the repository documentation. They have not been executed as a training run or validated on hardware for this article. Replace dataset and checkpoint paths with your own resources.
The retargeting stack recommends Ubuntu 20.04/22.04, Isaac Gym Preview 4, and Python 3.8. The bundled GR00T documentation recommends Python 3.10 and CUDA 12.4 for the VLA environment. Keep those environments separate rather than assuming one installation will support everything.
Download the archive from the official GitHub repository, extract it, and name the directory OpenTrajBooster. Then prepare the VLA environment:
cd OpenTrajBooster/VLA_model/gr00t_modified_for_OpenWBC
conda create -n trajbooster-vla python=3.10
conda activate trajbooster-vla
pip install --upgrade setuptools
pip install -e '.[base]'
pip install --no-build-isolation flash-attn==2.7.1.post4
python scripts/gr00t_finetune.py --help
Check the driver, PyTorch build, and CUDA environment before investigating FlashAttention compilation failures. If imports fail while displaying command help, solve that first instead of downloading a large dataset. The README documents a specific OpenCV CV_8U problem; apply its workaround when you encounter that error rather than reinstalling packages indiscriminately.
Starting from the released PPT checkpoint avoids repeating the entire intermediate training stage. The project also releases the Agibot-to-G1 retargeted dataset. Its README estimates roughly 30 GB for the dataset and 6 GB for the checkpoint. Actual downloads depend on the revision and files selected.
7. The data contract: 40 state dimensions, 32 actions
The converter's modality.json is essential reading before training:
| Group | Raw state dimensions | Action dimensions |
|---|---|---|
| Left and right arms | 14 | 14 |
| Left and right hands | 14 | 14 |
| Left and right legs | 12 | 0 |
| Base motion | 0 | 4 |
| Total | 40 | 32 |
The model observes leg states without directly predicting twelve leg action values. The final four action dimensions are worker commands, not four extra joint angles. Raw states can also undergo feature transforms before entering the network; the stored dataset dimension and the processed model feature dimension are different concepts.
The openwbc_g1 configuration expects three video streams: ego_view, wrist_left, and wrist_right. It uses the current observation and sixteen future actions. Confirm camera order, annotation names, and dimension ranges instead of relying on a generic LeRobot dataset loader to infer them.
The top-level README provides this three-view conversion command:
cd OpenTrajBooster/OpenWBC_to_Lerobot
python convert_3views_to_lerobot.py \
--input_dir /data/g1_raw \
--output_dir /data/g1_lerobot \
--dataset_name pick_toy \
--robot_type g1 \
--fps 30
The converter documentation also includes a single-view example using convert_to_lerobot.py. Do not silently replace the three-view script with that command. Likewise, the example conversion rate of 30 FPS and inference rate of 20 Hz require an explicit timing policy. They do not automatically represent the same action interval.
Before a long training run, inspect consecutive frames and actions from several episodes. Check whether the image timestamp matches the command, left and right hands are correctly assigned, and the language describes the recorded task. Split training and validation by episode so adjacent frames from one demonstration do not appear in both sets.
8. Training: PPT followed by target-robot fine-tuning
The study uses 1,960 episodes across 176 tasks, producing roughly 35 hours of retargeted data. PPT runs for 60,000 steps on two A100 80 GB GPUs with batch size 128. Target fine-tuning uses approximately ten minutes of data, comprising 28 episodes, with one A100, batch size sixteen, and 3,000 steps.
The crucial interpretation is that ten minutes refers to additional target-domain demonstrations. Pretrained weights, source demonstrations, retargeting, and GPU training still exist upstream. It is not the total data or compute budget of the project.
An example target fine-tuning command, adapted from the VLA README, starts from a locally downloaded PPT model:
cd OpenTrajBooster/VLA_model/gr00t_modified_for_OpenWBC
python scripts/gr00t_finetune.py \
--base-model-path /models/PPTmodel4UnitreeG1 \
--dataset-path /data/g1_lerobot \
--output-dir ./save/pick_toy_ppt \
--data-config openwbc_g1 \
--embodiment-tag new_embodiment \
--num-gpus 1 \
--batch-size 16 \
--max_steps 3000 \
--report-to tensorboard
Matching the number of steps does not reproduce the paper by itself. Dataset splits, camera setup, normalization, frozen modules, and code revision must also match. Begin with a short run that confirms all modalities load and checkpoints save before committing to the full schedule.
The two-stage procedure is worth testing against a controlled baseline. Keep your target demonstrations fixed, train one model with the PPT initialization and another without it, and evaluate both under the same conditions. Otherwise, additional data collection or a changed task can obscure whether transfer actually helped.
9. Results: separate complete and partial success

Table II evaluates each task over ten trials. The table below reports complete task success separately from partial outcomes:
| Task | PPT + 3K steps | No PPT + 3K steps | No PPT + 10K steps |
|---|---|---|---|
| Pick Mickey Mouse | 100% | 0% | 80% |
| Store Toys | 70% | 0% | 30% |
| Clean the Table | 70% | 0% | 70% |
| Pick Orange & Place | 10% | 0% | 0% |
The orange task exposes an important limitation. The PPT model has an additional 50% of trials with successful grasping but failed placement. Adding those to report 60% task success would misrepresent the result. With only ten trials, one outcome changes the reported rate by ten percentage points.
Table III moves the toy to positions absent from the teleoperation demonstrations. The PPT model reaches 80% success, while the no-PPT model trained for 10K steps reaches 0%. The authors' trajectory analysis suggests the baseline tends to repeat the demonstrated path, whereas the PPT model adapts its approach.

Loss curves are useful for tracking optimization, but they do not replace task evaluation. A model that closely predicts the demonstrated motion may fail when the object moves sideways. Include both task success and changed object positions in evaluation, and record where each attempt failed.
The “Pass the Water” demonstration is zero-shot with respect to target-robot teleoperation for that task. The task appears in PPT source data but not in the G1 fine-tuning set. It is not a skill absent from every stage of training.
10. Inference and deployment checks
Start with offline evaluation on separate episodes before connecting actuators:
python scripts/eval_policy.py \
--plot \
--model_path ./save/pick_toy_ppt/checkpoint-3000 \
--dataset-path /data/g1_validation \
--embodiment-tag new_embodiment \
--data-config openwbc_g1 \
--modality-keys base_motion left_hand right_hand
Look for abrupt base-command changes, unreasonable requested heights, switched hand labels, and isolated spikes. An average error can hide a short interval with poor behavior. Validation episodes should use the same schema as training while remaining independent demonstrations.
Real deployment additionally requires the image server, HomieDeploy, Unitree SDK2, and the correct network setup. After validating the low-level stack separately, the repository's VLA interface takes this form:
python scripts/G1_inference.py \
--arm=G1_29 \
--hand=dex3 \
--model-path ./save/pick_toy_ppt/checkpoint-3000 \
--goal pick_toy \
--frequency 20 \
--vis \
--filt
The README recommends an initial pose close to the teleoperation data and 20 Hz operation. Inspect outputs offline, verify the controller in simulation, and then conduct supervised hardware trials with an appropriate stopping mechanism. A filtering flag does not establish that every controller interaction is correct.
For a first experiment, use one object and two manipulation heights. Compare the same target dataset with and without PPT, then move the object to a new position. Record perception, grasping, placement, and body-stability failures separately. That breakdown tells you whether the next improvement belongs in the vision input, demonstration coverage, or worker controller.
TrajBooster offers a concrete recipe for turning another robot's data into a useful foundation for whole-body VLA. Its paper still identifies limitations in hand precision, visual consistency, available loco-manipulation data, and experiment scale. A narrow, measurable transfer experiment is a sensible starting point before attempting a household generalist.



