VnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
VnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam
VnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
  1. Home
  2. Blog
  3. FlowVLA-RL: GRPO for SmolVLA on LIBERO
wholebody-vlaflowvla-rlsmolvlagrpoliberolerobotflow-matchingvlarobot-rl

FlowVLA-RL: GRPO for SmolVLA on LIBERO

A practical FlowVLA-RL guide: turn SmolVLA's flow-matching sampler into an SDE with log-probs, then fine-tune online with GRPO on LIBERO.

Nguyễn Anh TuấnOctober 2, 202614 min read
FlowVLA-RL: GRPO for SmolVLA on LIBERO

FlowVLA-RL is one of the most useful open projects to study if you want to fine-tune a compact Vision-Language-Action model with online reinforcement learning without renting a cluster. The headline demo is simple: same LIBERO-Object task, same initial scene, the SFT policy on the left times out at the 280-step cap, while the GRPO-refined policy on the right finishes in 123 steps. The interesting part is not just the qualitative clip. The project addresses the exact mathematical gap that blocks ordinary policy-gradient RL on SmolVLA: SmolVLA's action expert is a deterministic flow-matching sampler, so there is no native log pi, no likelihood, and therefore no PPO-style ratio.

The original project is BlackMirean/FlowVLA-RL, with the detailed report FlowVLA-RL: Online GRPO for a Flow-Matching VLA on LIBERO. The SFT and GRPO checkpoints are published at MorpheusTzz/smolvla-grpo-libero-object. This is not a separate peer-reviewed paper yet; read it as a reproducible validation study with code, a technical report, public checkpoints, a locked evaluation protocol, and rollout videos. It builds on LeRobot's SmolVLA, LIBERO, Flow-GRPO's ODE-to-SDE idea, and the broader line of online RL for flow-based VLA policies such as piRL.

If you are new to VLA models, start with SmolVLA: Train a 450M VLA on a consumer GPU. If you already understand LIBERO and GRPO, NS-VLA v2: Fine-tune VLA on LIBERO is a useful comparison point. This article focuses on FlowVLA-RL itself: installation, training, inference, results, and the failure modes that can quietly invalidate your numbers.

Tool recommendations

VLA train/deploy stack

Train on cloud/workstation, then deploy optimized models to Jetson or the robot computer.

Cloud GPU for VLA / policy training Use for imitation learning, diffusion policies, RL, and robotics model fine-tuning. View cloud GPU → NVIDIA Jetson Orin NX / Orin Nano Edge deployment hardware for perception, logging, and optimized inference. View Jetson → Hugging Face / robotics dataset hosting Host datasets, checkpoints, and model cards for cleaner LeRobot/VLA workflows. View platform →

What Problem Does FlowVLA-RL Solve?

SmolVLA is a 450M-parameter VLA. It reads camera observations, a language instruction, and robot state, then emits a continuous action chunk. The key component is the action expert. Instead of autoregressively generating discrete action tokens, it starts from noise and integrates a learned velocity field for about 10 denoising steps to produce a 50-step action chunk. This is efficient for continuous control, but the default inference path is a deterministic ODE.

Deterministic sampling is fine for behavior cloning. You have demonstrations, you optimize a supervised flow-matching objective, and the model learns to map observations to action chunks. Online RL is different. PPO and GRPO need the probability of the sampled trajectory under the current policy and under the old or reference policy. If your sampler is only a deterministic integration path from noise to action, there is no tractable log-probability for the sampled action sequence.

FlowVLA-RL solves this by retrofitting the sampler:

Original SmolVLA:
noise -> deterministic ODE -> action chunk

FlowVLA-RL:
noise -> stochastic SDE transitions -> action chunk + log-prob

Once each denoising transition has a Gaussian log-density, the project can run critic-free GRPO. A group contains 4 rollouts from the same initial scene. The reward is the binary success signal from the official LIBERO checker. A rollout's advantage is reward - group_mean_reward. If two rollouts succeed and two fail, the successful ones get positive advantage and the failed ones get negative advantage. If all four succeed or all four fail, the group has zero advantage and contributes no useful gradient.

FlowVLA-RL learning curve with SDE training success and deterministic ODE evaluation probes — source: BlackMirean/FlowVLA-RL repo
FlowVLA-RL learning curve with SDE training success and deterministic ODE evaluation probes — source: BlackMirean/FlowVLA-RL repo

The plot is worth reading carefully. Training-time SDE success swings widely, while the deterministic ODE probes carry the main claim. The project locks the evaluation protocol before reporting headline numbers: deterministic ODE sampling, 20 episodes per task, n=200 per suite, evaluation seed 1000, and step caps of 280/280/300 for Object, Spatial, and Goal.

Architecture: SmolVLA Plus Flow-SDE GRPO

The code is easier to understand if you split it into three layers.

1. SmolVLA backbone. SmolVLA uses a compact VLM stack with a SigLIP vision encoder and a SmolVLM2/SmolLM-style language backbone, then attaches an action expert of roughly 100M parameters. The action expert receives VLM features, robot state, and the noise level, then predicts a velocity used to denoise the action chunk. In FlowVLA-RL, the VLM and vision encoder are frozen. The trainable portion is mostly the action expert plus state/action projections, about 99.9M parameters.

2. ODE-to-SDE retrofit. SmolVLA's original sampler is an ODE. FlowVLA-RL uses the predicted velocity v_theta(x, tau) to derive a score and a reverse-time drift:

score = grad log p_tau(x) = -(x + (1 - tau) * v_theta(x, tau)) / tau
reverse drift = v_theta - 0.5 * g_tau^2 * score
x_{tau - dt} ~ Normal(x - dt * reverse_drift, g_tau^2 * dt)

You do not need to memorize the math to run the repository, but you do need the intuition. Each step is now a Gaussian transition with a known mean and variance, so the log-probability can be summed over denoising steps, action dimensions, and executed chunk positions. LIBERO uses a 7-dimensional action space, while SmolVLA's action tensor has 32 dimensions, so FlowVLA-RL masks out the 25 padding dimensions when computing log-probs.

3. Episode-level GRPO. GRPO here uses no learned critic. For a shared initial scene, the sampler rolls out 4 stochastic SDE episodes. Final reward is 1 for LIBERO success and 0 for failure. The advantage is reward minus group mean, without standard-deviation normalization. The loss is a PPO clipped surrogate over log-density ratios, plus a k3 KL term to a frozen SFT reference with beta = 0.01. A critical runtime invariant protects the implementation: with epochs = 1, the first rescoring pass should reproduce rollout log-probs, so the ratio should be approximately 1. If that check fails, the training objective is not measuring the policy that collected the data.

Environment Setup

FlowVLA-RL pins the stack tightly. The report uses lerobot v0.6.0, Python 3.12.3, PyTorch 2.11 with CUDA 12.8, hf-libero 0.1.4, MuJoCo 3.8.1, and robosuite 1.4.0. The experiments ran on a 32GB RTX 5090, but measured SFT peak memory was 20.4 GiB, so the author argues that a 24GB GPU is enough to reproduce the work. Rollouts are CPU-bound: about 9.7 seconds per episode with only 8% GPU utilization.

A minimal setup following the README looks like this:

curl -LsSf https://astral.sh/uv/install.sh | sh
uv venv --python 3.12 .venv
source .venv/bin/activate

uv pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
git clone https://github.com/huggingface/lerobot.git
cd lerobot
git checkout v0.6.0
uv pip install -e ".[libero,smolvla]"
cd ..

git clone https://github.com/BlackMirean/FlowVLA-RL.git
cd FlowVLA-RL
uv pip install -e .
export MUJOCO_GL=egl

On a headless server, MUJOCO_GL=egl matters. Without a working EGL path, LIBERO may fail rendering or become far slower than expected. Before launching a long run, execute the unit tests and a smoke run:

python -m pytest tests/

python scripts/train_grpo_libero.py \
  --suite libero_object \
  --task-ids 0,1,2 \
  --groups 3 \
  --group-size 4 \
  --checkpoint <SFT_CKPT>

The smoke test is not meant to produce a good policy. It verifies the full path: load a checkpoint, create LIBERO environments, collect a group, compute log-probs, rescore, update, and save.

Step 1: Build or Download the SFT Baseline

FlowVLA-RL does not start RL directly from lerobot/smolvla_base. It first creates an SFT baseline on LIBERO. The report uses the full LIBERO dataset revision a1aaacb7..., trains for 25,000 steps with batch size 64, and fixes seed 1000. The command in the report is:

lerobot-train \
  --policy.path=lerobot/smolvla_base \
  --policy.device=cuda \
  --policy.push_to_hub=false \
  --policy.empty_cameras=1 \
  --rename_map='{"observation.images.image": "observation.images.camera1", "observation.images.image2": "observation.images.camera2"}' \
  --dataset.repo_id=lerobot/libero \
  --dataset.revision=a1aaacb7f6cd6ee5fb43120f673cebb0cfea7dd4 \
  --steps=25000 \
  --batch_size=64 \
  --save_freq=5000 \
  --seed=1000 \
  --output_dir=artifacts/sft_100pct \
  --job_name=sft_100pct

Beginners often miss rename_map. LIBERO observations in LeRobot do not use the exact camera keys expected by the SmolVLA policy config. If the mapping is wrong, the model may still run, but visual observations enter the wrong slots. policy.empty_cameras=1 is also not decorative. The checkpoints are in a SmolVLA format with three camera slots, while LIBERO provides two cameras, so the third slot is intentionally empty.

If you do not want to rerun SFT, download the public checkpoints:

hf download MorpheusTzz/smolvla-grpo-libero-object \
  --local-dir ckpts

The model repository contains sft-100pct-baseline/ and grpo-seed11-update300/. Under FlowVLA-RL's locked protocol, the SFT baseline scores 58.5% on LIBERO-Object, and the seed-11 GRPO checkpoint at update 300 scores 68.5%.

Step 2: Run Online GRPO on LIBERO-Object

The headline training run is LIBERO-Object from the 100% SFT checkpoint:

python scripts/train_grpo_libero.py \
  --checkpoint artifacts/sft_100pct/checkpoints/last/pretrained_model \
  --suite libero_object \
  --task-ids 0,1,2,3,4,5,6,7,8,9 \
  --group-size 4 \
  --groups 300 \
  --save-every 50 \
  --eta 0.4 \
  --lr 1e-6 \
  --kl-beta 0.01 \
  --seed 11

The important knobs are:

Argument Role Report value
group-size Rollouts from the same initial scene for group advantage 4
groups Number of usable GRPO updates to collect 300 for the headline
eta SDE exploration noise scale 0.4
lr Action expert learning rate 1e-6
kl-beta KL penalty to frozen SFT reference 0.01
seed Training seed, strongly affects rise timing 11, 31, 32...

The repository discards collapsed groups. That means the actual number of episodes is larger than groups * group_size. In the 400-update seed-11 run, the report needed 605 group attempts to obtain 400 usable groups, for 2,420 total episodes. The reason is that 33.9% of groups were all-success or all-fail, giving zero advantage.

Log three metric families. First, training success under the SDE sampler, but do not compare it directly to deterministic ODE evaluation. Second, ratio, clip fraction, and KL, which tell you whether the PPO-style objective is numerically sane. Third, collapsed group rate. If collapsed groups dominate, GRPO has little signal even though the simulator is busy.

FlowVLA-RL seed family: every seed improves, but the rise timing varies substantially — source: BlackMirean/FlowVLA-RL repo
FlowVLA-RL seed family: every seed improves, but the rise timing varies substantially — source: BlackMirean/FlowVLA-RL repo

Step 3: Inference and Evaluation

Evaluation uses deterministic ODE sampling, not the stochastic SDE sampler used during training. This is the most important protocol rule. If you train with SDE and evaluate with SDE, you may get a different number, but it is not comparable to the SFT baseline. FlowVLA-RL evaluates like this:

python scripts/run_baseline_eval.py \
  --checkpoint ckpts/grpo-seed11-update300 \
  --suite libero_object \
  --episodes 20 \
  --batch-size 5 \
  --seed 1000

Each suite has 10 tasks and 20 episodes per task, so n=200. With only 200 episodes, statistical noise is visible. The repository notes that reevaluating the same checkpoint can move results by roughly +/-2 percentage points, while a single-point comparison below about +/-4.9 pp is easy to overread. That is why the headline reports a fixed-budget mean across seeds and states its caveat.

There is also a LIBERO evaluation trap. In lerobot v0.6.0, the environment computes _max_episode_steps, but the environment loop does not enforce truncation the way you might expect. If a custom evaluator trusts the env to stop itself, rollouts may exceed the intended cap and inflate success. FlowVLA-RL enforces the step caps at the caller: 280 for Object, 280 for Spatial, and 300 for Goal.

Main Results

The headline result is GRPO on LIBERO-Object: SFT moves from 58.5% to a fixed-budget mean of 64.8%, a +6.3 percentage-point gain at 300 updates averaged across three seeds. All four probed seeds improve, but the rise timing varies. Seed 11 reaches 68.5% at update 300, +10.0 pp over baseline, then drops to 62.5% at update 400. That decline is why you should not report the best probe point casually.

Checkpoint LIBERO-Object success Note
SFT 100% baseline 58.5% n=200, deterministic ODE
GRPO seed 11 update 300 68.5% public checkpoint on Hugging Face
Fixed-budget 3-seed mean 64.8% +6.3 pp over baseline

The repository also measures SFT sample efficiency. The surprising finding is that reducing from 42.3 demonstrations per task to 4.7 demonstrations per task does not clearly hurt the three-suite mean, while dropping further to 1.9 demonstrations per task costs 9.0 pp. The author interprets this as SmolVLA saturating around five demonstrations per task on LIBERO under this setup.

SFT sample efficiency: the breakpoint sits between 1.9 and 4.7 demonstrations per task — source: BlackMirean/FlowVLA-RL repo
SFT sample efficiency: the breakpoint sits between 1.9 and 4.7 demonstrations per task — source: BlackMirean/FlowVLA-RL repo

Just as important, FlowVLA-RL does not work everywhere. The same recipe is nearly flat when started from the 3%-data checkpoint on Goal and Spatial: 16 probe points average only +1.0 pp. The report's explanation is concrete. Goal and Spatial are heavily scene-layout dominated. When four rollouts share one scene, 72-78% of groups become all-success or all-fail, so the advantage is zero. Episode-level binary-reward GRPO helps when the task is behavior-limited; it is much less helpful when outcomes are dominated by scene layout or by a memorized, rigid starting policy.

Debugging: Two Bugs That Make RL Look Fine While Doing Nothing

The first bug is a bf16/fp32 mismatch. Rollouts were collected under bf16 autocast, but rescored in fp32. Because a chunk log-prob sums thousands of Gaussian terms, tiny per-term errors accumulate into an order-one log-prob difference. The report shows ratios dropping near 0.34, outside the PPO clip range, so the objective loses the advantage signal. The fix is to use one shared autocast context for collection and rescoring, plus a runtime invariant: on the first epoch, the ratio should be close to 1.

The second bug is LIBERO scene reset. On Spatial and Goal, hard reset can resample object placement through robosuite's private RandomState, which means rollouts inside the same GRPO group are not actually starting from the same scene. Once group identity is broken, group-relative advantage no longer means what it should. FlowVLA-RL fixes this with snapshot-based soft reset: capture sim.get_state() after the first episode settles, then restore that state within the group. The acceptance check is strict: observations after group resets must match exactly.

These two cases are the practical lesson of the project. Robotics RL is not just an algorithm choice. It is a measurement system. If collection and rescoring disagree, or if environment resets violate the group invariant, training can run for hours while dashboards look superficially normal.

When Should You Use FlowVLA-RL?

Use FlowVLA-RL when you already have a healthy SmolVLA SFT checkpoint, a trustworthy success reward, and a benchmark or robot task with behavioral headroom. LIBERO-Object is a good example: the reward is sparse, but policies generate enough within-scene variation for GRPO to separate better and worse rollouts.

Be cautious if the policy is nearly random, if scene layout dominates success, or if you cannot guarantee group reset identity. On a real robot, the requirements are higher: safety constraints, reset automation, synchronized video/action/state logs, and a reliable stop mechanism. You should also avoid treating the 68.5% seed-11 number as a universal promise. It is a specific checkpoint, protocol, suite, and seed.

The reusable engineering pattern is the real value: freeze the VLM, train the smaller action expert, convert the flow sampler into a stochastic sampler with tractable log-probs, run group-relative online RL from shared initial scenes, and evaluate with a locked deterministic protocol. That is a practical path from imitation learning to online improvement for small VLA models.

Beginner Checklist

  1. Install the pinned stack: lerobot v0.6.0, Python 3.12, compatible CUDA PyTorch, and MUJOCO_GL=egl.
  2. Run python -m pytest tests/ to check SDE math, GRPO invariants, and rollout utilities.
  3. Download sft-100pct-baseline/ from Hugging Face or train the 25k-step SFT baseline.
  4. Evaluate the baseline with deterministic ODE, n=200, eval seed 1000, and correct step caps.
  5. Run the 3-group GRPO smoke test before launching a 300-group run.
  6. Watch first-pass ratio, KL, clip fraction, collapsed group rate, and rollout videos.
  7. Evaluate checkpoints at updates 50/100/200/300 under the same locked protocol.
  8. Compare only numbers with the same sampler, suite, cap, and evaluation seed.

Related Posts

  • SmolVLA: Train a 450M VLA on a Consumer GPU
  • NS-VLA v2: Fine-tune VLA on LIBERO
  • TGRPO: Fine-tune VLA with Trajectory GRPO
NT

Nguyễn Anh Tuấn

Robotics & AI Engineer. Building VnRobo — sharing knowledge about robot learning, VLA models, and automation.

Khám phá VnRobo

Fleet MonitoringROS 2 IntegrationAMR Solutions

Related Posts

Tutorial
X-VLA ICLR 2026: Soft-Prompted VLA 0.9B cho beginner LeRobot
x-vlavlaiclr-2026
wholebody-vla

X-VLA ICLR 2026: Soft-Prompted VLA 0.9B cho beginner LeRobot

Hướng dẫn X-VLA — flow-matching VLA 0.9B đạt SOTA trên 6 sim + 3 robot thật, native LeRobot, code open-source HuggingFace.

5/20/202611 min read
NT
Tutorial
NS-VLA v2: Fine-tune VLA trên LIBERO
ns-vlavlalibero
wholebody-vla

NS-VLA v2: Fine-tune VLA trên LIBERO

Hướng dẫn NS-VLA v2: neuro-symbolic primitives, BC warmup, GRPO/AWR fine-tuning và inference trên LIBERO.

9/7/202613 min read
NT
Tutorial
Guided Action Flow với SmolVLA
guided-action-flowsmolvlaq-guided-flow
wholebody-vla

Guided Action Flow với SmolVLA

Hướng dẫn GAF: dùng Q-guided critic để cải thiện SmolVLA trên LIBERO mà không fine-tune lại policy gốc.

9/2/202615 min read
NT
VnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam