Most VLA manipulation failures do not look dramatic. The robot reaches toward the object, pauses while the next policy call finishes, resumes a little late, closes the gripper a few frames off, then lifts an object that was never grasped cleanly. The high-level task was understood. The motion was almost right. The failure came from timing.
RACE-VLA is interesting because it attacks that exact timing problem. The paper, When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models, was submitted to arXiv on 2026-10-05 by Seonghoon Yu and collaborators from KAIST, GIST, NVIDIA, and POSTECH. The official repository is Seonghoon-Yu/RACE-VLA. At the time of writing, the repository contains the README and a real-robot GIF demo, while the code is marked as coming soon. So this guide does not pretend there is an official training command yet. Instead, it explains the method carefully enough that you can reproduce the idea on top of a flow-matching VLA such as pi0.5.

The Problem: Longer Chunks Reduce Pauses, But Transitions Break
Action chunking means the policy predicts several future actions per policy call instead of a single next action. A chunk length of 5 means the robot can execute five actions before it needs another policy output. A chunk length of 20 gives the robot a much larger execution buffer. This is attractive for VLA systems because large vision-language-action models are expensive to run. If inference takes longer than the current chunk takes to execute, the robot has no fresh command and must wait. That is the stop-and-go behavior RACE is designed to reduce.
The simple fix is to make chunks longer. Unfortunately, longer chunks are more open-loop. The robot has fewer chances to re-observe the scene, correct pose drift, or recover from a slightly wrong gripper contact. The key observation in RACE is that errors are not spread evenly across the chunk. They spike around subskill transitions: approach to grasp, grasp to lift, lift to place, or slide to insert.
These transition points are where the action distribution changes quickly. Before a grasp, the wrist may be moving toward the object. At the grasp, the gripper must close at the right moment. Right after the grasp, the arm should lift without dragging the object. A two-frame timing error can be enough to fail. This is why a model can look competent with short chunks and become unreliable when you stretch the horizon.

RACE reframes long-chunk execution as a “when to switch” problem. Instead of only asking the policy to predict future actions, it asks the action expert to predict where subskill transitions are likely to occur inside the chunk, then conditions action generation on that timing prior.
The Core Idea
A modern flow-matching VLA usually has two major parts. The first part is a vision-language model that encodes the current images and instruction into context features. The second part is an action expert that turns Gaussian noise into an action chunk through several denoising steps. In the paper, the action chunk is written as A_t = (a_t, ..., a_{t+H-1}), where H is the chunk length.
RACE adds three pieces around this process.
First, it runs an auxiliary one-step denoising pass from the initial noise. This pass is not used as the final action output. Its job is to expose action-token hidden states, one per future action position in the chunk. These hidden states are a rough motion sketch.
Second, it feeds those hidden states, together with cached VLM features, into a transition-timing prediction head. The head uses cross-attention from action tokens to VLM features, self-attention across the action tokens, a linear projection, and a sigmoid. The output is a vector p_hat of length H. Each element is a score for how close that timestep is to a subskill transition.
Third, RACE restarts full denoising from the same initial noise and injects the transition prior into every denoising step. The paper does this through adaptive RMSNorm-style modulation, controlled by a learnable per-step gate alpha^k. The gate lets the model decide how strongly the transition prior should influence each denoising step.
In plain language: RACE first makes a quick draft, uses that draft to estimate where the robot will need to switch subskills, then generates the real action chunk with that switch timing in mind.
Why No Manual Subskill Labels Are Needed
The method would be much less useful if every dataset needed human labels like approach, grasp, lift, and place. RACE avoids that. The paper derives transition labels from demonstration actions using PELT change-point detection. PELT looks for abrupt changes in the action trajectory, which often correspond to subskill boundaries.
After detecting change points, RACE converts them into soft labels. Instead of saying “this exact frame is the transition and all others are not,” it creates a peak around the transition and lets the score decay with distance. This matters because real manipulation transitions are rarely single-frame events. A grasp transition may span several control steps: wrist slows down, gripper starts closing, contact occurs, and lift begins.
For a beginner implementation, this is the simplest mental model:
for each demonstration episode:
read the action sequence A[0:T]
compute action deltas and smoothed velocities
run change-point detection on that signal
turn each change point into a soft peak
slice those soft peaks into labels for each training chunk
You should always visualize the labels. Plot joint motion, gripper command, and the detected transition peaks. If the peaks roughly land on approach-grasp-lift-place boundaries, the labels are useful. If they are firing on noise, timestamp glitches, or controller artifacts, fix the data before tuning the model.
A Practical Architecture Blueprint
Because the official code has not been released yet, here is a faithful implementation blueprint rather than a fake command-line recipe.
Images + instruction
|
v
Frozen VLM encoder
|
v
Cached visual-language features
|
+--> auxiliary one-step action expert
| |
| v
| action-token hidden states
| |
| v
| transition head -> p_hat[0:H]
|
v
full flow-matching denoising conditioned on p_hat
|
v
long action chunk
|
v
asynchronous execution on the robot
Your base VLA needs to expose the action expert hidden states per action position. If your model already uses a DiT-like or transformer action head, this is straightforward. If your model has a simpler MLP policy, you can still borrow the idea, but the transition head will be less naturally aligned to action tokens.
The transition head should be small compared with the VLM. In the paper, the VLM is frozen and cached for both passes. Only the action expert and the added RACE modules are trained. This keeps the adaptation lightweight and makes RACE closer to post-training than full VLA retraining.
Dataset Preparation
A useful RACE dataset contains at least:
| Field | Purpose |
|---|---|
front_image, wrist_image |
visual observations for the VLM |
instruction |
language command for the task |
robot_state |
joint or end-effector state |
action |
target robot action at each timestep |
episode_id, timestamp |
required for chunk slicing and transition detection |
The real-robot experiment in the paper uses 30 demonstrations of a pick-and-place task: “Pick up the gray bowl next to the plate and place it on the plate.” The robot setup is an AgileX PiPER arm, a 6-DoF lightweight arm with 626 mm reach and a 1.5 kg payload, equipped with a gripper. The observations come from a wrist camera and a front camera. The demonstrations record images, robot state, and actions at 30 Hz. The action is the absolute position of six joints plus the gripper.
For your first reproduction, resist the temptation to start with a broad multi-task benchmark. Use one real task with clean demonstrations. Vary the object pose enough to test generalization, but not so much that perception and control failures become impossible to separate. RACE is about long-chunk reliability; it is not a replacement for camera calibration, timestamp alignment, or good teleoperation data.
Training Step
Each RACE training step has two passes and three losses.
The auxiliary pass starts from noise and runs one unmodulated denoising step. It gives hidden states and a velocity prediction. The transition head predicts p_hat, supervised by the soft transition label p_t. This produces L_timing, typically binary cross-entropy. The auxiliary velocity also receives a flow-matching loss, L_aux, so the hidden states remain motion-aware.
The full pass samples a flow time s, mixes noise and ground-truth action chunk accordingly, and trains the action expert to predict the velocity field while conditioned on transition timing. During training, the paper uses teacher forcing: it conditions on the target transition label rather than the predicted prior, but jitters the label by plus or minus one step with probability 0.5. This makes the action expert robust to small timing errors at inference.
The total loss is:
L = L_full + lambda_aux * L_aux + lambda_timing * L_timing
For the real-robot setup, the paper fine-tunes from the pretrained pi0.5 base model for 20K steps on 30 demonstrations. It uses AdamW, gradient clipping at 1.0, batch size 32, and a cosine learning-rate schedule with 1K warmup, decaying from 2.5e-5 to 2.5e-6, without EMA. The RACE model is trained with chunk length H=40.
Inference Step
At inference, RACE adds only one auxiliary pass and one small head before the full denoising loop.
1. Encode the current images and instruction with the VLM.
2. Sample the initial action noise A0.
3. Run one auxiliary denoising pass without transition modulation.
4. Predict p_hat[0:H] with the transition head.
5. Restart full denoising from A0.
6. At denoising step k, inject alpha^k * p_hat.
7. Return the final action chunk.
8. Execute the chunk asynchronously while requesting the next chunk.
The real-robot policy predicts a chunk twice as long as it executes. For example, with H=40, the robot executes H_exec=20 actions, then asks for the next chunk while the remaining 20 actions provide a buffer for inference. In the repository demo, the baseline replans every 5 actions while RACE replans every 20 actions. That is the 4x action-chunk extension in the title.
This is why RACE complements real-time chunking instead of replacing it. Asynchronous execution hides inference latency by generating the next chunk during the current one. RACE makes the current chunk long and reliable enough that the robot does not run out of commands before the next chunk arrives.
Results That Matter
On VLABench, the pi0.5 baseline at H_exec=5 reports an average success rate of 24.6 and a progress score of 40.8. RACE at the same execution length improves to 30.5 success rate and 45.8 progress score. That means the transition-aware objective helps even before using longer chunks.
The more important test is long execution. With H_exec=20, RACE reports 33.0 average success rate and 47.8 progress score in the main VLABench table. The paper also trains three random seeds for that setting and reports a mean of 33.3 ± 0.5 success rate and 48.6 ± 0.9 progress score. That is a useful sign that the result is not a single lucky run.
On the real robot, the numbers are easier to interpret:
| Method | H_exec |
Success rate | Completion time | Idle time |
|---|---|---|---|---|
pi0.5 fine-tuning |
5 | 36% | 12.6 s | 2.11 s |
pi0.5 fine-tuning |
20 | 48% | 11.6 s | 0.40 s |
| RACE | 20 | 66% | 11.8 s | 0.41 s |
The jump from H_exec=5 to H_exec=20 reduces idle time from 2.11 seconds to about 0.41 seconds, roughly a 5x reduction. Fine-tuning with the long chunk already helps idle time, but its success rate is only 48%. RACE keeps the same low idle time while raising success to 66%. That is the core result: long chunks solve the waiting problem, and transition conditioning makes long chunks safer.

Common Implementation Mistakes
The first mistake is treating transition prediction as a separate classifier that does not backpropagate into the action expert. In RACE, L_timing also updates the action expert through the auxiliary pass. That encourages the action hidden states themselves to carry transition-aware information.
The second mistake is using hard one-hot transition labels. Real subskill switches are fuzzy. Soft labels reduce brittleness and match the physical reality of manipulation.
The third mistake is deploying a long chunk on hardware before measuring action errors in replay. Start with offline replay and simulation. Compare short-chunk baseline, long-chunk fine-tuning, and RACE. Look specifically at action error around detected transition points. If RACE does not reduce those spikes offline, it is not ready for the robot.
The fourth mistake is ignoring safety. Longer chunks mean longer open-loop execution. Add joint-limit checks, action-delta clamps, stale-camera detection, and an emergency replan path. RACE makes long chunks more reliable, not magically safe under every disturbance.
When To Use RACE
RACE is a strong fit for tasks with clear phase changes: pick-and-place, drawer opening, object insertion, stacking, simple tool use, and mobile manipulation with approach-contact-transfer stages. It is especially useful when short-horizon VLA control works, but the robot becomes stop-and-go or unreliable when you try to increase the execution horizon.
It is less useful if your main bottleneck is perception failure, poor calibration, missing gripper feedback, or a dataset with inconsistent teleoperation. In those cases, transition timing is not the first problem to fix.
The broader lesson is elegant. Many efficient VLA methods try to make each policy call cheaper. RACE reduces the number of calls by making each call cover more useful time. For real robots, that distinction matters. A smoother robot is not just a faster model; it is a policy whose timing matches the physical subskills of the task.



