VnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
VnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam
VnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
  1. Home
  2. Blog
  3. Offline Supervision VLA-RL: 50% Cheaper RL
wholebody-vlavlareinforcement-learningopenvlaoffline-supervisionmanipulationppo

Offline Supervision VLA-RL: 50% Cheaper RL

A practical guide to RefKL and DataBC: use offline demos as a PPO prior to cut online RL steps for OpenVLA manipulation.

Nguyễn Anh TuấnSeptember 30, 202614 min read
Offline Supervision VLA-RL: 50% Cheaper RL

Offline Supervision VLA-RL: using demos to cut online RL cost by 50%

Online reinforcement learning is one of the most powerful ways to improve a robot manipulation policy, but it is also the expensive part of the VLA training stack. PPO needs environment rollouts, reward computation, episode resets, trajectory storage, value estimation, and repeated policy updates. In simulation this means GPU time and wall-clock time. On real robots it also means hardware wear, human supervision, manual resets, and collision risk.

The ICML 2026 workshop paper "Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models" by Dmitriy Poyarkov, Aleksei Staroverov, and Aleksandr I. Panov asks a practical question: if we already have offline demonstrations for supervised fine-tuning, can we keep using that signal during online RL instead of discarding it after initialization?

The answer is yes. On the RL4VLA benchmark with OpenVLA LoRA adaptation, two guided PPO variants, RefKL and DataBC, reach performance close to standard PPO trained for 2M environment steps while using only 1M steps. The key idea is that offline supervision should not be treated only as a warm start. It can act as an optimization prior inside PPO during the early, noisy phase of online learning.

If you have read our OpenVLA deep dive, VLA-RL scaling overview, or SimpleVLA-RL training guide, this article is the next layer: how to combine supervised demonstrations and online RL in one objective so manipulation policies train faster without giving up OOD robustness.

Tool recommendations

VLA train/deploy stack

Train on cloud/workstation, then deploy optimized models to Jetson or the robot computer.

Cloud GPU for VLA / policy training Use for imitation learning, diffusion policies, RL, and robotics model fine-tuning. View cloud GPU → NVIDIA Jetson Orin NX / Orin Nano Edge deployment hardware for perception, logging, and optimized inference. View Jetson → Hugging Face / robotics dataset hosting Host datasets, checkpoints, and model cards for cleaner LeRobot/VLA workflows. View platform →

The beginner version

Imagine teaching a robot to place an object on a plate. You usually have two learning signals:

Signal Strength Weakness
Offline demonstrations Cheap, stable, easy to collect with teleoperation or motion planning The policy can fail when it drifts away from the dataset distribution
Online RL Improves through closed-loop feedback and reward Expensive, slow at the beginning, and sensitive to sparse rewards

Offline-only SFT is like asking the robot to watch examples and imitate them. Online RL is like asking the robot to practice by itself. Offline Supervision VLA-RL combines the two: while the robot practices with PPO, it is still softly guided by the old demonstrations or by a frozen SFT policy trained on those demonstrations.

The paper studies three ways to combine the signals:

  1. PPO SFT-init: train an SFT LoRA first, then initialize PPO from that checkpoint.
  2. RefKL: train PPO while adding a KL penalty toward a frozen SFT reference policy.
  3. DataBC: train PPO while adding a behavior cloning loss on offline demonstration batches.

The important result is that PPO SFT-init helps, but it is not enough. RefKL and DataBC are stronger because the offline signal remains active during RL optimization. In other words, the demonstrations continue to shape learning instead of merely defining the starting point.

Original paper and repository

The core sources are:

  • Project page: alstar8.github.io/offline-supervision-vla-rl
  • arXiv: arXiv:2607.19399
  • GitHub repository: alstar8/offline-supervision-vla-rl

The repository implements hybrid offline-online RL fine-tuning of OpenVLA LoRA adapters on RL4VLA. It includes ManiSkill, SimplerEnv, openvla, real2sim, sim2real, and experiment launch scripts under scripts. The structure is typical for modern VLA experiments: model code and SFT live around openvla, environments and RL loops live around ManiSkill/SimplerEnv, and paper experiments are wrapped by shell scripts.

OpenVLA architecture with value head for PPO - source: alstar8/offline-supervision-vla-rl repo
OpenVLA architecture with value head for PPO - source: alstar8/offline-supervision-vla-rl repo

Task setup: RL for VLA manipulation

The paper models control as a partially observable Markov decision process. At each timestep, the policy receives:

  • the current RGB observation I_t;
  • a natural-language instruction l, such as "put the carrot on the plate";
  • no direct access to the full simulator state.

The policy predicts tokenized actions u_t. The environment wrapper converts those tokens into executable robot commands a_t. Episodes have a fixed horizon of T = 80, and the reward is a sparse difference-based signal derived from task progress. This is not a dense shaping reward where every tiny motion is manually scored.

The backbone is OpenVLA-warmup. The base model is frozen, and training updates a rank-32 LoRA adapter. For RL, the system uses the RL4VLA value-head modification so PPO can estimate advantages. The main task is PutOnPlateInScene25Main-v3: a simulated robot must place an object on a plate.

For beginners, the most important detail is that this is not full fine-tuning of a 7B model. It is parameter-efficient adaptation. That matters because full fine-tuning would change the compute cost, stability properties, and catastrophic forgetting behavior.

Main idea: offline supervision as a PPO prior

Standard PPO updates a policy using the likelihood ratio between the current policy and the old rollout policy:

r_t(theta) = pi_theta(a_t | o_t) / pi_theta_old(a_t | o_t)

The clipped PPO objective prevents updates from becoming too large. The problem is that robot manipulation rewards are often sparse. Early PPO can be almost blind: the policy has not learned useful action structure yet, successful episodes are rare, and exploration is inefficient.

Offline Supervision VLA-RL adds an auxiliary loss to PPO:

L(theta) = L_PPO(theta) + beta * L_aux(theta)

The auxiliary loss can be defined in two ways:

  • L_RefKL: keep the current policy close to a frozen SFT reference policy.
  • L_DataBC: keep the current policy close to offline demonstration actions through behavior cloning.

The coefficient beta is scheduled. The paper holds beta = beta_0 for the first 100k steps, linearly decays it to zero over the next 200k steps, and then continues as pure PPO after 300k steps. Intuitively, the demonstrations guide the early phase, then the constraint fades so RL can optimize the final policy.

PPO, RefKL, and DataBC training pipelines - source: alstar8/offline-supervision-vla-rl repo
PPO, RefKL, and DataBC training pipelines - source: alstar8/offline-supervision-vla-rl repo

How RefKL works

RefKL uses an SFT policy as a frozen reference model. During RL, the current policy is penalized if its action distribution moves too far away from that reference:

L_RefKL = average KL(pi_ref(. | o) || pi_theta(. | o))

The full objective becomes PPO plus the weighted KL term. If the current policy starts drifting away from useful demonstration behavior, the KL penalty pulls it back. But the penalty is soft, so PPO can still move the policy when online reward shows that a different behavior is better.

The advantage of RefKL is that it uses the full action distribution of the reference policy. For a large VLA, that distribution can contain richer information than a single demonstrated action. A demonstration may show one valid motion, but the reference policy may encode a smoother set of nearby preferences.

The paper finds RefKL to be the most consistent guided variant. It stays above standard PPO across training curves and is generally more reliable than DataBC. This makes sense: KL-to-reference is a smooth regularizer, while BC on a fixed offline dataset can over-anchor the policy to states seen in the dataset.

How DataBC differs

DataBC does not need a reference-policy forward pass for every offline supervision signal. It directly adds a behavior cloning loss:

L_DataBC = - E log pi_theta(a_star | o)

Here (o, a_star) comes from the offline demonstration dataset. This is simple and easy to implement if your training stack already has an offline dataloader. It is also close to traditional supervised fine-tuning.

The downside is that the dataset only contains states visited by the demonstrator or motion planner. During online RL, the policy will visit new states, especially after mistakes. In those states, a fixed dataset BC term may be less informative and may sometimes pull the policy back toward the old distribution.

DataBC still beats PPO at the same 1M-step budget, but it is not as consistently strong as RefKL. If you need a quick baseline and have limited memory, DataBC is a good starting point. If you have a strong SFT checkpoint and enough GPU memory to evaluate it, RefKL is the stronger default.

Installation

The commands below follow the official repository README. The environment assumes CUDA 12.1 and PyTorch 2.2. You will need a Linux machine with a sufficiently large NVIDIA GPU for OpenVLA 7B LoRA work, or a multi-GPU setup if you plan to reproduce multiple seeds.

git clone https://github.com/alstar8/offline-supervision-vla-rl.git
cd offline-supervision-vla-rl

# Fast path: create the rlvla-guided environment
bash ./install_env.sh --flash-attn-wheel /path/to/flash_attn.whl
conda activate rlvla-guided

Manual installation:

conda create -n rlvla-guided python=3.10 -y
conda activate rlvla-guided

pip install torch==2.2.0 torchvision==0.17.0 torchaudio==2.2.0 \
  --index-url https://download.pytorch.org/whl/cu121

cd openvla && pip install -e . && cd ..
pip install -U tyro datasets==3.3.2
cd ManiSkill && pip install -e . && cd ..
cd real2sim && pip install -e . && cd ..
cd SimplerEnv && pip install -e . && cd ..

The RLDS dataset builder uses a separate environment:

cd openvla/rlds_dataset_builder
conda env create -f environment_ubuntu.yml
conda activate rlds_env

For beginners, the likely failure points are CUDA/PyTorch compatibility, flash-attn, and editable installs across submodules. Debug layer by layer: first import torch, then mani_skill, then SimplerEnv, and only then launch training.

Preparing demonstration data

The experiments use PutOnPlateInScene25Main-v3. The repository collects demonstrations with the ManiSkill motion planner:

conda activate rlvla-guided
cd ManiSkill

python -m mani_skill.examples.motionplanning.widowx.collect_simpler \
  -e "PutOnPlateInScene25Main-v3" \
  --save_video \
  --save_data \
  --num_procs 16 \
  --num_traj 16400 \
  --seed=100

Then build the SFT RLDS dataset:

conda activate rlds_env
cd openvla/rlds_dataset_builder/sft_dataset

tfds build --overwrite
mkdir -p ../../../datasets
mv -T ~/tensorflow_datasets/example_dataset ../../../datasets/sft

The paper uses an early-stopped SFT reference checkpoint trained on 2k demonstrations for 7.5k steps. That detail is useful: the reference policy does not need to be a huge, fully converged SFT model. A reasonably good supervised checkpoint can already provide a valuable RL prior.

Training: SFT, PPO, RefKL, DataBC

The main scripts are:

Method Script
SFT reference scripts/train_sft.sh
Standard PPO scripts/train_ppo.sh
PPO SFT-init scripts/train_ppo_sft_init.sh
RefKL scripts/train_refkl.sh
DataBC scripts/train_databc.sh

Train the SFT reference:

bash scripts/train_sft.sh

Run the PPO baseline:

bash scripts/train_ppo.sh --seed 0

Run PPO initialized from the SFT LoRA:

bash scripts/train_ppo_sft_init.sh --seed 0

Run the two guided methods:

# RefKL: PPO + KL penalty to frozen SFT reference
bash scripts/train_refkl.sh --seed 0

# DataBC: PPO + offline behavior cloning auxiliary loss
bash scripts/train_databc.sh --seed 0

To cap training at 1M environment steps:

bash scripts/train_refkl.sh --seed 0 --steps_max=1000000

According to the repository mapping, RefKL corresponds to train_ms3_ppo_bc_teacher.py with --bc_to_ref_enabled, DataBC corresponds to train_ms3_ppo_sft.py, and the beta curriculum uses flags such as --bc_to_ref_coef, --bc_to_ref_hold_steps=100000, and --bc_to_ref_decay_steps=300000.

Inference and evaluation

In this paper, "inference" mainly means evaluating a trained checkpoint inside RL4VLA/SimplerEnv. It is not a direct real-robot deployment recipe. Set the checkpoint path and run:

CKPT_PATH="gen-robot/openvla-7b-rlvla-warmup" \
UNNORM_KEY="bridge_orig" \
VLA_LOAD_PATH="../SimplerEnv/wandb/<run>/glob/steps_XXXXXX" \
bash scripts/eval_policy.sh

Then aggregate success rates:

cd SimplerEnv/scripts
python calc_statistics.py

The evaluation protocol is:

  • 64 in-distribution episodes per checkpoint.
  • 960 out-of-distribution episodes in total.
  • 15 OOD environments, 64 episodes each.
  • OOD settings grouped into action, language, and vision shifts.
  • Evaluation seeds {0, 1, 2}.

RL4VLA in-distribution and OOD scenes - source: alstar8/offline-supervision-vla-rl repo
RL4VLA in-distribution and OOD scenes - source: alstar8/offline-supervision-vla-rl repo

Key results

The main comparison:

Method IND OOD Act OOD Lang OOD Vis OOD Avg
SFT 0.82 0.46 0.60 0.74 0.62
PPO SFT-init (1M) 0.87 0.51 0.69 0.78 0.69
PPO (2M) 0.92 0.82 0.75 0.76 0.77
PPO RefKL (1M) 0.93 0.79 0.76 0.76 0.77
PPO DataBC (1M) 0.91 0.74 0.73 0.74 0.74

The headline result is simple: PPO RefKL at 1M steps reaches an OOD average of 0.77, matching PPO at 2M steps. Its IND score is also slightly higher, 0.93 versus 0.92. DataBC is also strong: 0.74 OOD average at 1M steps.

At the same 1M budget:

Method IND OOD Avg
PPO (1M) 0.75 0.64
PPO SFT-init (1M) 0.87 0.69
PPO RefKL (1M) 0.93 0.77
PPO DataBC (1M) 0.91 0.74

RefKL improves OOD average by +0.13 over PPO at 1M steps. DataBC improves it by +0.10. For robotics, those gains are substantial because every point of success rate usually costs many more rollouts.

OOD learning curves for RefKL and DataBC versus PPO - source: alstar8/offline-supervision-vla-rl repo
OOD learning curves for RefKL and DataBC versus PPO - source: alstar8/offline-supervision-vla-rl repo

Why not just initialize from SFT?

The obvious baseline is reasonable: train SFT, initialize PPO from the SFT LoRA, and continue with online RL. The paper shows that this is better than plain PPO at 1M steps, but it is still weaker than RefKL and DataBC.

One interpretation is that after the LoRA adapter has fit the supervised distribution strongly, online RL has a harder time changing the action behavior, especially under action OOD shifts. The policy starts from a better point, but its later adaptation can be slow.

RefKL is different. It does not only start from SFT. It injects supervised knowledge into the objective while keeping PPO active. Because beta decays to zero, the policy is guided early but still becomes a pure RL policy later. For deployment teams, the lesson is important: an SFT checkpoint is not only an initialization; it can also be a teacher during RL.

When should you use this?

Use Offline Supervision VLA-RL when:

  • you already have demonstration data from teleoperation or motion planning;
  • your task has sparse rewards, such as pick-place, insertion, drawer opening, or tool use;
  • PPO or GRPO is too slow during the first 20-30% of training;
  • you adapt a large VLA such as OpenVLA with LoRA or another PEFT method;
  • you care about OOD generalization, not only one in-distribution scene.

Be careful when:

  • demonstrations are low-quality or mismatched with the online task;
  • the reward encourages shortcuts;
  • you do not have enough memory to run a frozen reference policy;
  • your environment reset or evaluation protocol is very different from RL4VLA.

If memory is tight, start with DataBC. If you can afford the reference forward pass, use RefKL as the stronger default.

Reproduction checklist

A practical checklist:

  1. Clone the repository and create rlvla-guided.
  2. Create the separate rlds_env for dataset building.
  3. Collect demonstrations for PutOnPlateInScene25Main-v3.
  4. Build the RLDS SFT dataset.
  5. Train the SFT reference or use the checkpoint path expected by the scripts.
  6. Run PPO for 1M steps to establish your local baseline.
  7. Run RefKL for 1M steps with beta curriculum.
  8. Run DataBC for 1M steps as a cheaper guided baseline.
  9. Evaluate IND and all 15 OOD environments with the same seeds.
  10. Compare OOD average, not just IND success.

During early debugging, you can reduce seeds to save time. For a report or internal benchmark, keep the evaluation protocol close to the paper. If you only look at one in-distribution scene, you may miss the entire point of the method.

Connection to LeRobot and real robots

This method is not a drop-in LeRobot recipe yet, but the idea transfers cleanly. In a LeRobot-style workflow, you often have:

  • an offline dataset collected by teleoperation;
  • an SFT or behavior cloning policy;
  • a simulator or real-robot online learning loop;
  • a reward from success detection, a scripted metric, human intervention, or a learned evaluator.

Instead of throwing away the dataset after SFT, keep it active during online RL. If you have a frozen SFT policy, add KL-to-reference like RefKL. If not, sample batches from the dataset and add a BC loss like DataBC. Our HIL-SERL with LeRobot guide is a good companion for the online RL and human-intervention side; Offline Supervision VLA-RL adds the idea that demonstration data should stay alive inside the objective.

Conclusion

Offline Supervision VLA-RL matters because it targets the core pain point of robotics: online RL is powerful but expensive. The paper shows that offline demonstrations should not be used only once for SFT. When offline supervision is added to PPO through RefKL or DataBC, the policy learns faster, keeps OOD generalization, and RefKL can match the OOD average of 2M-step PPO with only 1M environment steps.

For beginners, the big lesson is to treat a demonstration dataset as a continuing training prior, not a file you use once and discard. For VLA manipulation teams, this is a practical route to reduce compute and interaction cost before moving toward real robot training.

Related Posts

  • OpenVLA: open VLA for robots
  • VLA-RL: online RL for VLA manipulation
  • SimpleVLA-RL (3): Setup & Training
NT

Nguyễn Anh Tuấn

Robotics & AI Engineer. Building VnRobo — sharing knowledge about robot learning, VLA models, and automation.

Khám phá VnRobo

Fleet MonitoringROS 2 IntegrationAMR Solutions

Related Posts

Research
TGRPO: Fine-tune VLA với Trajectory GRPO và LLM Dense Reward
vlagrporeinforcement-learning
wholebody-vla

TGRPO: Fine-tune VLA với Trajectory GRPO và LLM Dense Reward

TGRPO kết hợp LLM dense reward với dual-level advantage (step + trajectory) để fine-tune OpenVLA-7B trên LIBERO, đạt 91% success rate — vượt SFT 4.6% và PPO 4.4%.

7/1/202611 min read
NT
Tutorial
LeRobot v0.6: Reward Models và lerobot-eval CLI
lerobotreinforcement-learningreward-model
wholebody-vla

LeRobot v0.6: Reward Models và lerobot-eval CLI

Hướng dẫn dùng Robometer, TOPReward, lerobot-eval CLI với 6 benchmarks và DAgger để đóng vòng RL tự động cho manipulation VLA trong LeRobot v0.6.

7/22/202611 min read
NT
Research
FORCE: Tăng 79% success rate khi fine-tune VLA bằng RL
vlareinforcement-learningfine-tuning
wholebody-vla

FORCE: Tăng 79% success rate khi fine-tune VLA bằng RL

FORCE giải quyết 2 điểm yếu cốt lõi của RL fine-tuning VLA — Q-function không ổn định và data exploration kém chất lượng — qua Value-Calibrated Warm-Up và Self-Distillation, đạt 79% cải thiện tuyệt đối mà không cần human intervention.

7/3/202611 min read
NT
VnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam