When a humanoid falls, is motion tracking enough?
Imagine teleoperating a humanoid while standing upright. Someone pushes the robot onto the floor, but your reference pose remains upright. A controller trained mainly to follow clean demonstrations may keep minimizing joint error even though the robot has entered a very different physical situation. The useful question becomes: which states can this body move through to stand up again?
UFO, from the RoboParty Lab Team, investigates that question through unsupervised reinforcement learning. Its policy learns a latent space of reusable whole-body behavior. Motion tracking and goal reaching then become different ways to request behavior from the same motor policy. This tutorial focuses on temporal reachability and recovery, complementing the tracking architecture covered in GR00T SONIC architecture.
The primary sources are the UFO technical report, updated July 14, 2026, the project page, and Roboparty/UFO on GitHub. The report presents UFO-FB, based on Forward-Backward representation learning, and UFO-TeCH, which learns temporal-distance representations. All numerical results discussed here are author-reported; VnRobo has not independently reproduced these experiments.
Where WBC, WBID/QP, and RL fit
Whole-body control coordinates the legs, torso, and arms so the robot can accomplish a target while maintaining feasible motion. Reaching forward changes the body's balance requirements. The legs cannot simply ignore what the arms are doing, even when each actuator has its own local controller.
In whole-body inverse dynamics, or WBID, a controller commonly solves for joint accelerations, torques, and contact forces. A quadratic program, or QP, can minimize tracking error while enforcing dynamics and constraints such as torque limits and friction. This approach depends on a model, a state estimator, and assumptions about contact. It also needs an appropriate reference or planner: asking for a standing pose does not automatically specify a complete get-up sequence.
RL learns an observation-to-action mapping by interacting with simulation. Tracking RL typically rewards following a reference motion. UFO instead learns reusable representations and a latent-conditioned actor, allowing the current body state to influence how a requested behavior is reached.
| Approach | Typical control request | What must be prepared |
|---|---|---|
| WBID/QP | Task accelerations and contact specifications | Dynamics model, estimator, planner |
| Tracking RL | Reference pose or motion window | Retargeted data, rewards, simulation training |
| UFO-FB / UFO-TeCH | Observations plus a latent skill or goal | Motion prior, representation learning, off-policy RL |
UFO is a motor-control research framework rather than a language-conditioned VLA. A teleoperation interface or higher-level planner can supply its targets. Connecting a VLA still requires a defined interface from semantic decisions to motion targets or latent commands.
Unsupervised does not mean no data or no rewards
Here, unsupervised refers to learning reusable behavior without designing a separate downstream task reward for every skill. The framework still uses motion demonstrations, simulation objectives, and substantial engineering.
The policy explores the simulator and stores transitions in a replay buffer. Off-policy learning reuses previously collected experience instead of optimizing exclusively from the newest rollout. Motion demonstrations provide a behavioral prior, helping exploration remain connected to useful human-like movement.
There are also style and auxiliary objectives. The FB preset includes components related to action rate, joint limits, undesired contacts, and foot slippage. Some torque-related coefficients are zero in the default preset. Reading a list of penalty names is therefore insufficient to determine which constraints actively influence a particular run.
For a beginner, the practical consequence is simple: removing the motion prior and auxiliary objectives while keeping a goal-progress reward creates a different experiment. It does not reproduce the complete system evaluated in the report.
Architecture: observations, latent commands, and joint targets
The report's G1 setup has 29 degrees of freedom. The actor produces targets for joint PD controllers, rather than solving an explicit contact-force optimization at every control step. Its basic proprioceptive observation contains joint positions relative to a nominal pose, joint velocities, root angular velocity, and projected gravity. The report describes this vector as 64-dimensional.
Observation and action history help the actor infer dynamics that are difficult to recover from a single measurement. In the implementation, the actor uses the input groups state, last_action, and history_actor. Training networks can additionally access privileged_state, containing information available in simulation. Separating these inputs matters because a deployable actor cannot rely on arbitrary simulator-only measurements.
The inspected FB and TeCH presets use a 256-dimensional latent. Their actor configuration is residual, with width 2048 and six hidden layers. These are preset values, not a promise about every downloadable checkpoint. Exported input dimensions should always come from the checkpoint and its metadata.
Retargeted motion data Robot observations + history
| |
Representation encoder |
| |
Latent command z ------------------> Actor
|
Joint PD targets
|
Humanoid G1
Training: simulator -> replay -> representation / style / auxiliary critics
With FB, the backward map embeds target states, while the forward map represents future-state occupancy under a latent-conditioned policy. This factorization supports reusing a learned representation for different downstream requests. UFO's --agent fb selects an implementation with style and auxiliary components as well; it is not merely a minimal textbook FB learner.
TeCH: learn distance through temporal reachability
Two poses can look similar in joint space while requiring very different physical transitions. A robot lying on the floor must establish contacts and shift its body before standing. A small pose error during a fall can still be difficult to correct. TeCH learns a representation intended to capture this temporal structure.

The encoder sees a current state, a consecutive state, and a pseudo-goal sampled from a future part of a trajectory. A contrastive objective distinguishes temporally distant states. A local regularization term keeps consecutive representations sufficiently continuous. The diagram illustrates reachability in a learned behavior space, rather than physical distance across a room.
The policy receives a progress reward: moving closer to the latent goal earns a positive signal, while moving farther away earns a negative one. A conceptual rendering of Equation 6 is:
r_progress = distance(goal_latent, encoded_current_state)
- distance(goal_latent, encoded_next_state)
This explanation is not a replacement training implementation. The full actor objective combines critics for representation progress, motion style, and auxiliary constraints. TeCH also differs from FB in an important way: the report says it does not provide the same general successor-feature-based reward-inference interface. For a first TeCH experiment, focus on tracking and goal reaching, rather than assuming arbitrary reward optimization is supported.
Installation: use the supported G1 path first
The following commands are intended for a Linux research workstation. They are documented reader instructions and were not executed during preparation of this article. The repository's pyproject requires Python 3.10, pins MJLab 1.4.0, uses MuJoCo around version 3.8, and constrains PyTorch to version 2.7 through below 2.8. Training also needs a compatible CUDA stack.
git clone https://github.com/Roboparty/UFO.git
cd UFO
python3 -m pip install --user uv
export PATH="$HOME/.local/bin:$PATH"
uv sync
bash scripts/download_data.sh g1_lafan
ls -lh humanoidverse/data/lafan_29dof_10s-clipped.pkl
The download script verifies the dataset checksum. This is processed motion data for G1, not an automatic human-to-any-robot retargeting service. If the expected file is missing, resolve that before launching training. A data-path problem will remain a data-path problem regardless of GPU count.
Run the documented smoke test next:
./run_train.sh \
--agent fb \
--data-manifest configs/data/example_mix.yaml \
--gpu-ids single \
--smoke \
--work-dir /tmp/ufo_smoke_g1
Check whether environments initialize, motion data loads, reported losses remain finite, and outputs are created. A smoke test exercises a short training loop. It does not establish that the policy has learned stable walking or that it is suitable for a physical robot.
Training: understand the resource budget
The published FB quick start includes a multi-GPU configuration:
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
./run_train.sh \
--agent fb \
--gpu-ids all \
--num-envs 1024 \
--num-env-steps 192000000 \
--work-dir runs/ufo_fb_g1 \
--data-path humanoidverse/data/lafan_29dof_10s-clipped.pkl \
--update-z-every-step 100 \
--buffer-size 5120000
num-envs and buffer-size are per GPU; num-env-steps is a global sample budget. Eight GPUs with 1024 environments each therefore means 8192 environments in total. Replay storage can consume substantial memory, so an eight-H200 setting should not be treated as an optimal configuration for a single consumer card.
To experiment with TeCH, use a separate work directory, change the agent to --agent tech, and use --update-z-every-step 10, following the quick start. Historical tldr names remain in implementation files and as a deprecated alias. Keep the two experiments' checkpoint directories separate so evaluation cannot accidentally load the wrong model.
The update-z interval controls how long a latent is held during rollouts. The report observes a tradeoff: more frequent updates can improve tracking accuracy, while less compressed behavior can become less smooth and more difficult to deploy. This setting is not the control frequency. The report's evaluation uses 200 Hz simulation and 50 Hz control.
For an affordable first experiment, keep the robot, data, and algorithm fixed while reducing the sample budget or environment count to fit available hardware. Record the GPU model, GPU count, dataset revision, seed, and wall-clock duration. Only then compare algorithms under matched evaluation conditions. Otherwise, a faster run might simply reflect a different workload.
Adding a skill without overwhelming the motion distribution
The report examines cartwheel injection as an example of adding a rare agile skill. If all trajectories are merged into one pool with prioritized sampling, difficult motions can repeatedly receive high priority. The learner then spends disproportionate effort on a narrow subset of behaviors, potentially destabilizing actor-critic optimization.
UFO instead chooses the dataset using a fixed probability first, then applies prioritized sampling within that source. The reported foundation/cartwheel mixture is 0.95/0.05. In those experiments, naive merged-data baselines developed non-finite metrics around 14–28 million environment steps; fixed-ratio mixing remained stable to 384 million steps. That is evidence for the tested setup, not a universal prescription that every new skill should receive exactly five percent of samples.
The current configs/data/example_mix.yaml contains only LaFAN with weight 1.0. Adding another source requires valid, already retargeted motion data and an intentional source weight. Supported import formats include RobotState CSV/NPZ and ufo_pkl.
For another robot, a MuJoCo XML, correct joint ordering, and reviewed robot semantics are essential. The configuration generator produces drafts, not validated actuator gains or contact definitions. Shared infrastructure also does not imply a shared checkpoint: changing action or observation dimensions requires a compatible trained model.
Inference: evaluate a full motion and export ONNX
Tracking inference should use full motion sequences, rather than only the clipped training file. Replace the placeholder below with an actual compatible full-motion dataset:
CUDA_VISIBLE_DEVICES=0 \
uv run python -m humanoidverse.tracking_inference \
--model-folder runs/ufo_fb_g1 \
--data-path /path/to/full_motions.pkl \
--device cuda:0 \
--headless \
--save-mp4 \
--motion-list 0 \
--export-onnx true
Outputs go under the model folder's tracking_inference/ directory. Inspect clips for foot slipping, root-heading errors, and discontinuities. Keep metrics alongside videos so a successful highlight cannot hide failures elsewhere in the trajectory.
Export metadata records robot configuration, XML path, controlled joints, actor input dimensions, latent dimension, and action dimension. A 29-element output does not by itself prove correct joint ordering. Matching normalization, history layout, and action scaling is equally necessary.
For goal reaching, an encoder maps a target state into a latent goal. For motion tracking, the report aggregates encoded future states across a look-ahead window. This creates an important distinction between offline motion playback and live teleoperation: a recorded trajectory provides future references that an online pose stream does not yet contain. Use the appropriate online runtime rather than assuming those input conditions are interchangeable.
The training and inference guide also documents goal and reward modules. Their existence does not establish support for every algorithm and robot combination. Non-G1 reward semantics are more limited, and TeCH has the representation-related limitations described earlier.
Teleop-to-policy: use a compatible deployment runtime
The main branch contains training, import, inference, and export. The separate deploy branch provides a G1 29-DoF runtime. Its released artifact includes the policy ONNX, a backward encoder ONNX, and latent context files. The README identifies xuewang/ufo-g1-policy as the model artifact source.
The online pipeline can receive PICO/XRobot motion, retarget it with code vendored in the deployment runtime, encode the target into z, and transmit that latent through ZeroMQ. The actor closes the loop using robot proprioception. The operator requests movement; the motor policy computes joint actions from the body's current state.
Start with sim2sim after installing the separate deployment environment, downloading the complete artifact, and checking its release manifest. The documented two-terminal example is:
# Terminal A: UFO-Deploy directory, ufo-deploy environment
python -m sim_env.base_sim \
--robot_config ./config/robot/g1.yaml \
--scene_config ./config/scene/g1_29dof.yaml
# Terminal B: same directory and environment
python rl_policy/ufo_policy.py \
--robot_config config/robot/g1.yaml \
--policy_config config/policy/g1_policy.yaml \
--model_path model/g1_policy/exported/FBcprAuxModel.onnx \
--task config/exp/tracking/tracking.yaml
The simulator controls documented in the README include i for standing-pose interpolation, ] to enable the policy, and [ to start tracking. Physical deployment additionally requires the matching ARM64 control binding, the correct DDS interface, and functioning stale-latent and stop-latch handling. Fast workstation inference alone does not validate that deployment path.
In this article, teleop-to-policy means feeding online motion targets into a pretrained motor policy. It does not mean the repo automatically records demonstrations and trains a complete VLA. For downstream imitation learning, log synchronized observations, references, latents, and robot actions, then define which policy layer is being learned.
Results: read hardware and evaluation protocol together

The highlight figure advertises a 6–8 hour training region. Table 1 of the report gives the more specific hardware comparison:
| Method | Hardware | Environments | Training time |
|---|---|---|---|
| BFM-Zero | 1 H200 | 1024 | 45 hours |
| UFO-FB | 1 H200 | 1024 | 21 hours |
| UFO-FB | 8 H200 | 8 × 1024 | 6 hours |
| UFO-FB | 8 RTX 4090 | 8 × 512 | 14 hours |
These are wall-clock times to the authors' deployable-policy criterion. They are not smoke-test durations or single-GPU promises. The report contains some inconsistent short descriptions of the RTX 4090 duration; use its explicit table row for budgeting, and measure your own configuration.
Table 3 reports joint-position tracking error, retaining the authors' metric labels and values:
| Method | LAFAN1 | 100-Style |
|---|---|---|
| SONIC, full trajectory | 0.2916 ± 0.1887 | 0.1522 ± 0.1049 |
| SONIC, before termination | 0.1081 ± 0.0213 | 0.1355 ± 0.0351 |
| BFM-Zero, full trajectory | 0.1510 ± 0.0255 | 0.1674 ± 0.0459 |
| UFO-FB, full trajectory | 0.1178 ± 0.0575 | 0.1278 ± 0.0403 |
| UFO-TeCH, full trajectory | 0.1272 ± 0.0582 | 0.1385 ± 0.0392 |
Before-termination evaluation excludes the portion after failure. Full-trajectory evaluation includes what happens when a robot falls and tries to recover. These answer different questions, so the table should not become a blanket claim that UFO wins every tracking scenario. TeCH also does not have the lowest tracking error here, despite the report describing advantages in discontinuous goal transitions and recovery.
The reported training corpus is approximately 2–3 hours of LAFAN1; the 100-Style test corpus is approximately 18–19 hours. Data duration is useful context, but it does not prove that your desired manipulation task, contact pattern, or operator motion is covered.
The authors demonstrate real-world pushes, pulls, falls, and get-up behavior. They also acknowledge weaker locomotion, sideways and backward walking issues, limited global tracking, and latency from backward mapping. The demonstrations establish that real-robot trials exist. They do not provide a standardized success rate across arbitrary environments or objects.
A first experiment that teaches you something
Choose a standing–squatting–standing motion in simulation. Keep the reference fixed, log joint error and foot contacts, and measure how long the robot takes to resume tracking after a controlled perturbation. Then try a small set of discontinuous goal poses and inspect the transitions it selects instead of supplying a complete recorded trajectory.
If joint tracking looks good while the robot drifts, separate root error from joint error. If a clip fails, preserve the recovery segment in your logs. When FB and TeCH behave differently, inspect normalization, observation history, and latent update cadence before attributing every difference to the learning objective.
UFO provides a useful platform for studying reusable motor skills and comparing FB with temporal-distance learning. Choosing a controller still depends on the actual goal: precise reference following, rapid pose transitions, stable teleoperation, or integration with a higher-level planner.



