VnRoboVnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
VnRobo logoVnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam
VnRoboVnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
  1. Home
  2. Blog
  3. RoboDojo: Benchmark π0.5, SmolVLA and GR00T
wholebody-vlavlarobodojoxpolicylabpi0smolvlagrootmanipulationbenchmark

RoboDojo: Benchmark π0.5, SmolVLA and GR00T

Set up RoboDojo and XPolicyLab, train and evaluate VLA policies across 42 simulation tasks, and interpret scores without misleading comparisons.

Nguyễn Anh TuấnOctober 11, 202614 min read
RoboDojo: Benchmark π0.5, SmolVLA and GR00T

Why a successful grasping demo is not enough

A robot receives an instruction, picks up a cup, and hangs it on a rack. The demonstration looks convincing. But what happens after the lighting changes, the target disappears under a cover, or the instruction requires several consecutive operations? Choosing a Vision-Language-Action policy requires repeatable measurements, visible failures, and a clear description of what was trained.

RoboDojo provides 42 simulation manipulation tasks and 18 physical tasks. XPolicyLab connects different policies to the evaluation environment while preserving their individual dependency stacks. This tutorial explains how to build a local evaluation workflow for π0.5, SmolVLA, and GR00T-N1.7, the particular GR00T version represented in the results discussed here.

Commands were checked against public documentation and scripts on October 11, 2026. Published results below belong to the benchmark authors; VnRobo has not independently executed these experiments. One distinction matters immediately: the paper labels SmolVLA Single Task. That row must not be described as a multitask checkpoint trained under identical conditions to the other models.

1. The research idea: measure distinct capabilities

A benchmark that only rearranges objects on a table can hide weaknesses in memory, contact control, and task composition. RoboDojo organizes evaluation around five capabilities. The official simulation task catalog explains their instructions and task conditions.

Dimension Tasks What it tests Examples
Generalization 12 Whether manipulation survives scene changes stack_bowls, fold_clothes
Memory 6 Whether earlier observations guide later actions cover_blocks, swap_T
Precision 8 Accurate placement and contact interaction insert_tubes, deposit_coin
Long-Horizon 8 Successful composition of multiple operations organize_table, fill_egg_holder
Open 8 Tasks without task-specific training demonstrations solve_equation, pour_by_language

Consider cover_blocks: after covering colored blocks, the robot must act using their previously observed colors. The current image alone cannot supply all the information. In insert_tubes, identifying the tube is only the beginning; orientation and final placement must match the rack. In organize_table, an early mistake can invalidate later steps even if each individual skill sometimes works.

Open tasks are evaluation-only in the benchmark's task data split. Adding demonstrations from those exact tasks to fine-tuning would change the experiment. Keep the intended split intact and document additional data rather than calling that modified setup open generalization.

RoboDojo manipulation tasks grouped by capability — source: RoboDojo paper
RoboDojo manipulation tasks grouped by capability — source: RoboDojo paper

Image source: the paper's simulation task overview. The featured image uses a separate project teaser from the RoboDojo README.

Tool recommendations

VLA train/deploy stack

Train on cloud/workstation, then deploy optimized models to Jetson or the robot computer.

Cloud GPU for VLA / policy training Use for imitation learning, diffusion policies, RL, and robotics model fine-tuning. View cloud GPU → NVIDIA Jetson Orin NX / Orin Nano Edge deployment hardware for perception, logging, and optimized inference. View Jetson → Hugging Face / robotics dataset hosting Host datasets, checkpoints, and model cards for cleaner LeRobot/VLA workflows. View platform →

2. Architecture: separate simulation from policy inference

The simulator creates the robot, objects, cameras, and task conditions. The policy server runs the model. The environment sends images, robot state, and an instruction; the adapter converts them into the model's inputs and returns an action chunk. The simulator executes actions and obtains new observations.

text
Isaac Sim / Isaac Lab
  scene + robot + cameras + task success checker
                 |
       observation / instruction
                 v
XPolicyLab client <--- WebSocket ---> policy server
                                      |
                               model adapter
                                      |
                         π0.5 / SmolVLA / GR00T
                                      |
                               action chunk
                 <--------------------+

XPolicyLab defines shared methods including update_obs, get_action, their batched counterparts, and reset. Resetting clears policy state between episodes. A memory policy that retains the previous episode's information can produce invalid measurements even when the simulator correctly resets its scene.

Environment isolation also prevents dependency conflicts. The π0.5 adapter uses uv-managed OpenPI, SmolVLA uses a LeRobot conda environment, and GR00T-N1.7 has its own uv stack. Installing every model into the Isaac Sim environment is likely to create version conflicts and makes failures harder to diagnose.

XPolicyLab infrastructure connecting policy servers to simulators and robots — source: XPolicyLab/XPolicyLab repository
XPolicyLab infrastructure connecting policy servers to simulators and robots — source: XPolicyLab/XPolicyLab repository

Image source: assets/infra.png in XPolicyLab.

What distinguishes the three policies?

π0.5 combines visual and language understanding with continuous action generation. The public OpenPI implementation supports the flow matching head for both training and inference: conditioned on observations, the model learns to transform noisy actions into executable action sequences. This is different from merely generating a textual plan. See the official OpenPI repository.

SmolVLA targets a smaller computational footprint with a compact vision-language backbone and an action expert. Its efficiency mechanisms include reducing visual tokens, and its broader deployment design supports asynchronous inference. Check the actual benchmark adapter before claiming a particular run enables asynchronous execution. A model feature and an evaluation setting are separate things. The Hugging Face architecture explanation provides the original design context.

GR00T-N1.7 is NVIDIA's foundation policy version used here. The examined integration defaults to nvidia/GR00T-N1.7-3B and includes a separate Cosmos configuration. Its modality configuration and embodiment_tag determine how images, state, and actions correspond to the benchmark robot. Results for N1, N1.5, and N1.7 should remain distinct. See the GR00T_N17 adapter.

3. Prepare a machine and install the simulator

The RoboDojo installation guide recommends Ubuntu 22.04 x64, at least 32 GB RAM, an NVIDIA GPU with at least 16 GB VRAM, Linux driver 570 or 580, and CUDA 12.8. These describe the simulation environment; they do not guarantee enough memory for full fine-tuning of every VLA model.

For context, OpenPI's single-GPU estimates exceed 22.5 GB for LoRA fine-tuning and 70 GB for full fine-tuning. Running simulation and inference on the same GPU adds their memory demands. Start with one episode, inspect memory consumption, and only then increase concurrency or choose a larger training configuration.

On your benchmark workstation, use a dedicated working directory:

bash
sudo apt install libvulkan1 mesa-vulkan-drivers vulkan-tools git-lfs
vulkaninfo | head
git lfs install
git clone https://github.com/RoboDojo-Benchmark/RoboDojo.git
cd RoboDojo
bash scripts/install.sh -i
conda activate RoboDojo
bash scripts/init_assets.sh
python utils/update_embodiment_config_path.py
bash scripts/robodojo.sh doctor
bash scripts/robodojo.sh dimensions

The installer creates the environment and installs dependencies. The asset script retrieves the simulation assets. The Python utility updates robot configuration paths to the local asset location. Skipping that last step can produce missing USD errors even after a successful download. Repeat the path update if you move the repository.

The current installer updates submodules from their remotes. Record the actual RoboDojo and XPolicyLab revisions after installation, alongside dataset and checkpoint versions. The README reports September 16–17 fixes for observation alignment and RGB channel ordering and requires XPolicyLab commit bb9a0b5 or later for the corresponding update. Mixing new code with old data can make an integration problem look like poor model performance.

4. Choose data and verify action alignment

Downloading the entire dataset is unnecessary for an initial wiring check. The downloader lists a demo bundle of roughly 1.5 GB, joint-only LeRobot v3.0 data of roughly 120 GB, and full simulation HDF5 data of roughly 523 GB. The depth export is approximately 4.5 TB, which is an expensive starting point for RGB evaluation. These estimates come from the official data download script.

bash
# From the RoboDojo root; select data for your intended workflow.
bash scripts/RoboDojo/download_data.sh huggingface demo

# Prepared joint-only data for compatible training workflows:
bash scripts/RoboDojo/download_data.sh huggingface lerobot_v3.0

The demo is a small bundle, not a promise that every training task is present. Inspect its contents before using it as a source. The π0.5 conversion example below requires raw HDF5 demonstrations for its selected task. Downloading a prepared LeRobot export does not automatically satisfy every adapter's raw-data conversion path.

Official LeRobot camera keys are observation.images.cam_high, observation.images.cam_left_wrist, and observation.images.cam_right_wrist. SmolVLA renames them internally at load time, so preserve the official on-disk keys. State and action fields must also match the selected control mode.

In simulation HDF5, state[t] represents the current state and action[t] is an absolute target for the next frame. Shared joint and gripper fields follow action[t] = state[t+1], with final-frame padding. Do not subtract state again if the adapter expects an absolute joint target. End-effector poses have a documented quaternion convention and coordinate frame; conventions from another dataset are not interchangeable.

Before training, watch the three camera previews, read the instruction, and inspect several state/action rows. If the left arm visibly moves while the right arm's joint values change, resolve the mapping first. A short inspection can prevent a long training run from learning a systematically incorrect relationship.

5. Training workflows and the published protocol

There are two useful starting points: download a benchmark checkpoint to validate evaluation, or fine-tune a foundation checkpoint using demonstrations. The first route separates deployment debugging from training expenditure.

bash
# From the root; the corresponding policy adapters must already exist.
bash scripts/RoboDojo/download_ckpt.sh huggingface Pi_05
bash scripts/RoboDojo/download_ckpt.sh huggingface SmolVLA
bash scripts/RoboDojo/download_ckpt.sh huggingface GR00T_N17

The downloader checks policy names and remote folders. If the selected mirror lacks a checkpoint, it fails rather than silently producing usable weights. Inspect each adapter's checkpoints directory after downloading and select the actual run folder and training step.

π0.5: convert within the OpenPI environment

The Pi_05 README documents HDF5-to-LeRobot conversion and source task selection. This is a single-task learning exercise, requiring raw stack_bowls demonstrations beforehand:

bash
cd XPolicyLab/policy/Pi_05
bash install.sh
bash process_data.sh RoboDojo stack_bowls arx_x5 joint
bash train.sh RoboDojo stack_bowls arx_x5 joint 0 0

The conventional run name is RoboDojo-stack_bowls-arx_x5-joint-0. The final argument selects the training GPU; the preceding zero is the seed. This exercise does not reproduce the paper's multitask π0.5 row. Multitask training requires the intended demonstration set, an appropriate dataset repo ID, and compatible values for OPENPI_TRAIN_CONFIG_NAME and OPENPI_LEROBOT_REPO_ID.

SmolVLA: use the LeRobot v3.0 dataset

The SmolVLA README has no top-level process_data.sh. Its training script consumes a dataset repo ID, mapping a task name to identifiers such as RoboDojo_sim_build_tower_v30. Make the corresponding export available through HF_LEROBOT_HOME, or set SMOVLA_REPO_ID to your real dataset.

bash
cd XPolicyLab/policy/SmolVLA
bash install.sh
# Run only after the repo ID and cache contain the correct build_tower data.
bash train.sh RoboDojo build_tower arx_x5 joint 0 0

LeRobot checkpoints include nested training-step artifacts. Check checkpoint_num in deploy.yml: evaluating a different step changes the model even when the outer run name remains identical. Keep that distinction in experiment records.

GR00T-N1.7: conversion and modality configuration

The GR00T adapter requires GR00T_LEROBOT_HOME to point to the dataset parent directory. Its documented processing path copies a prepared v3.0 export and downgrades it to v2.1. For arx_x5, the default source dataset is RoboDojo_sim_arx-x5_v30.

bash
cd XPolicyLab/policy/GR00T_N17
bash install.sh
export GR00T_LEROBOT_HOME=/absolute/path/to/lerobot/datasets
bash process_data.sh RoboDojo cotrain arx_x5 joint
bash train.sh RoboDojo cotrain arx_x5 joint 0 0

Replace the illustrative absolute path with your actual location. The converter accepts expert_data_num for compatibility but does not use it to subset episodes. A 50-demonstration ablation needs an actual subset dataset supplied through GR00T_SRC_DATASET. Naming a run “50ep” does not reduce its training data.

The paper reports the following simulation recipes, separately from the small exercises above:

Policy Batch size Training steps Important detail
π0.5 256 60,000 Fine-tuned from pi05_base
SmolVLA Single Task 512 100,000 Preserve the single-task designation
GR00T-N1.7 640 100,000 Match embodiment and modalities

Source: RoboDojo paper, Appendix K. Different batch sizes, optimization budgets, and task scopes mean this is not a controlled architecture-only comparison. Changing to LoRA or a smaller batch is legitimate experimentation, but it creates a different recipe that should be reported explicitly.

6. Inference: one episode before the full benchmark

First inspect deploy.yml: checkpoint selection, camera transforms, action mode, and environment must agree. π0.5 accepts uv for its policy environment; SmolVLA normally uses the smolvla conda environment; GR00T can also use uv. The examples use joint, requiring compatible absolute joint actions in data and checkpoints.

The following evaluates the single-task π0.5 checkpoint trained above. For downloaded weights, substitute the actual run name and confirm its deployment settings.

bash
# Return to the RoboDojo root before running this command.
export EVAL_ENV_TYPE=sim
bash scripts/robodojo.sh eval \
  --policy-dir XPolicyLab/policy/Pi_05 \
  --task stack_bowls \
  --ckpt RoboDojo-stack_bowls-arx_x5-joint-0 \
  --policy-env uv --eval-env RoboDojo \
  --env-cfg arx_x5 --action-type joint \
  --seed 0 --policy-gpu 0 --env-gpu 0 \
  --eval-num 1 --dry-run

The dry run prints dispatch arguments without running inference. Remove --dry-run after checking them. Watch the episode video: do the correct arm and gripper move, and are the cameras correct? A completed process demonstrates functioning execution; it does not necessarily demonstrate successful manipulation.

After a smoke test, select a checkpoint appropriate for the entire benchmark. This next command is a template: replace the uppercase directory and checkpoint values with your actual policy and run.

bash
bash scripts/robodojo.sh benchmark \
  --policy-dir XPolicyLab/policy/POLICY_DIRECTORY \
  --ckpt ACTUAL_CHECKPOINT_NAME \
  --policy-env uv --eval-env RoboDojo \
  --env-cfg arx_x5 --action-type joint \
  --seed 0 --eval-num native
bash scripts/robodojo.sh summarize

Use --policy-env smolvla when evaluating the default SmolVLA conda setup. native uses per-task episode counts, whereas one or five episodes are debugging settings. Inspect the runnable inventory with dimensions: Generalization's _random variants can make the dispatch list longer than the 42 conceptual tasks.

Different RoboDojo scenes evaluated concurrently — source: RoboDojo paper
Different RoboDojo scenes evaluated concurrently — source: RoboDojo paper

Image source: the paper's heterogeneous parallel simulation figure.

Parallel simulation reduces waiting time but consumes additional GPU memory. Begin sequentially, then use the multi-GPU options described in Quick Evaluation. Extra CPU cores alone do not guarantee useful throughput when rendering and policy inference compete for the GPU.

7. Results: distinguish score from success rate

This table reproduces the paper snapshot dated July 3, 2026, not the live ranking on this article's publication date. Each cell is score / success rate. Score captures task-defined progress, while success rate measures complete task success.

Policy Generalization Precision Long-Horizon Memory Open Average
π0.5 13.37 / 8.17% 12.40 / 5.50% 23.54 / 14.67% 5.78 / 4.56% 1.98 / 1.67% 11.41 / 6.91%
GR00T-N1.7 2.16 / 1.22% 2.54 / 0.67% 8.30 / 3.58% 1.06 / 0.89% 0.18 / 0.17% 2.85 / 1.31%
SmolVLA Single Task 1.69 / 1.22% 2.87 / 0.33% 1.22 / 0.25% 3.35 / 2.44% 0.00 / 0.00% 1.83 / 0.85%

Source: RoboDojo paper, Table 1. New entries are published on the official leaderboard. Keep their evaluation dates and recipes separate from this snapshot.

Average is the mean of the five capability dimensions. It is not simply total successes divided by every episode across all tasks. Because dimensions contain different numbers of tasks, those calculations have different weights and can give different results.

The paper describes three training seeds for most policies, with 50 trials per task for each seed. Generalization splits those trials into 25 standard and 25 randomized runs. Distinguish training seeds from evaluation layout seeds: three layouts evaluated with one checkpoint do not reproduce three independently trained policies.

π0.5 has the strongest Average among these three published rows, but a 6.91% success rate still indicates substantial limitations. It does not establish reliable deployment readiness. SmolVLA's Memory success rate exceeds GR00T's in this snapshot, while its Precision success rate is lower. Such differences motivate task-level investigation instead of a universal claim that one model dominates another in every situation.

8. Troubleshooting and fair comparisons

If every task fails immediately, investigate integration first. Use doctor for missing assets, examine policy-server logs for loading failures, verify the checkpoint step, inspect RGB previews, and confirm joint versus end-effector control and gripper ordering. Use the official JPEG decoder instead of adding an extra BGR/RGB conversion based on an OpenCV habit.

If early operations succeed but later ones fail, inspect action chunks and observation timing. Longer chunks may reduce inference calls while leaving the robot without fresh feedback for longer. That is a testable hypothesis, not a diagnosis established by one video. Keep the checkpoint fixed, change one setting, and rerun the same layout set.

A useful report records task coverage, asset and dataset versions, camera mapping, action mode, training recipe, checkpoint step, seeds, and trial counts. Then report scores, success rates, variability, latency, and VRAM consumption. Preserve failed rollouts alongside successful ones. Otherwise, a claim such as “π0.5 wins” cannot distinguish model effects from data or evaluation configuration.

For official publication, consult the current submission protocol. Local measurements are internal results until the project's verification process is satisfied. Simulation also cannot substitute for physical testing with the intended robot, camera geometry, and contact conditions.

For a beginner, the first milestone is one explainable episode with correct observations and actions. Then finish one task, evaluate one dimension, and expand to all 42 tasks under the intended protocol. Each increase in GPU expenditure should answer a concrete technical question, rather than simply producing a larger table.

Related Posts

  • Fine-tuning SmolVLA with LeRobot
  • GR00T-N1.7 data and fine-tuning
  • Long-horizon manipulation with VLA and LeRobot
NT

Nguyễn Anh Tuấn

Robotics & AI Engineer. Building VnRobo — sharing knowledge about robot learning, VLA models, and automation.

Explore VnRobo

Fleet MonitoringROS 2 IntegrationAMR Solutions
← Previous
Fine-Tune UnifoLM-WLA-1.0 for Unitree G1

Related Posts

Tutorial
FluxVLA Engine: Hands-On Guide to the One-Stop VLA Platform
vlalerobotpi0
wholebody-vla

FluxVLA Engine: Hands-On Guide to the One-Stop VLA Platform

LimX Dynamics' open-source FluxVLA Engine covers the full pipeline from data to real-robot deployment, supporting Pi0, GR00T, and OpenVLA in one unified framework.

9/17/20269 min read
NT
Tutorial
VLASH: Real-Time VLAs via Async Inference (11.8× Faster)
vlareal-timeasynchronous-inference
wholebody-vla

VLASH: Real-Time VLAs via Async Inference (11.8× Faster)

MIT Han Lab's VLASH makes VLAs real-time via future-state-aware asynchronous inference — 11.8× lower reaction latency, no architecture changes, open-source.

8/20/202612 min read
NT
Tutorial
FM-VLA: Force Memory Tokens for Contact-Rich Manipulation
vlaforce-sensingmanipulation
wholebody-vla

FM-VLA: Force Memory Tokens for Contact-Rich Manipulation

FM-VLA compresses force/torque history into compact tokens via VAE, letting VLAs count contacts and track interaction progress — 83.3% success on AgiBot G1.

7/31/202612 min read
NT
VnRobo logoVnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam