Offline Supervision VLA-RL: using demos to cut online RL cost by 50%
Online reinforcement learning is one of the most powerful ways to improve a robot manipulation policy, but it is also the expensive part of the VLA training stack. PPO needs environment rollouts, reward computation, episode resets, trajectory storage, value estimation, and repeated policy updates. In simulation this means GPU time and wall-clock time. On real robots it also means hardware wear, human supervision, manual resets, and collision risk.
The ICML 2026 workshop paper "Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models" by Dmitriy Poyarkov, Aleksei Staroverov, and Aleksandr I. Panov asks a practical question: if we already have offline demonstrations for supervised fine-tuning, can we keep using that signal during online RL instead of discarding it after initialization?
The answer is yes. On the RL4VLA benchmark with OpenVLA LoRA adaptation, two guided PPO variants, RefKL and DataBC, reach performance close to standard PPO trained for 2M environment steps while using only 1M steps. The key idea is that offline supervision should not be treated only as a warm start. It can act as an optimization prior inside PPO during the early, noisy phase of online learning.
If you have read our OpenVLA deep dive, VLA-RL scaling overview, or SimpleVLA-RL training guide, this article is the next layer: how to combine supervised demonstrations and online RL in one objective so manipulation policies train faster without giving up OOD robustness.
The beginner version
Imagine teaching a robot to place an object on a plate. You usually have two learning signals:
| Signal | Strength | Weakness |
|---|---|---|
| Offline demonstrations | Cheap, stable, easy to collect with teleoperation or motion planning | The policy can fail when it drifts away from the dataset distribution |
| Online RL | Improves through closed-loop feedback and reward | Expensive, slow at the beginning, and sensitive to sparse rewards |
Offline-only SFT is like asking the robot to watch examples and imitate them. Online RL is like asking the robot to practice by itself. Offline Supervision VLA-RL combines the two: while the robot practices with PPO, it is still softly guided by the old demonstrations or by a frozen SFT policy trained on those demonstrations.
The paper studies three ways to combine the signals:
- PPO SFT-init: train an SFT LoRA first, then initialize PPO from that checkpoint.
- RefKL: train PPO while adding a KL penalty toward a frozen SFT reference policy.
- DataBC: train PPO while adding a behavior cloning loss on offline demonstration batches.
The important result is that PPO SFT-init helps, but it is not enough. RefKL and DataBC are stronger because the offline signal remains active during RL optimization. In other words, the demonstrations continue to shape learning instead of merely defining the starting point.
Original paper and repository
The core sources are:
- Project page: alstar8.github.io/offline-supervision-vla-rl
- arXiv: arXiv:2607.19399
- GitHub repository: alstar8/offline-supervision-vla-rl
The repository implements hybrid offline-online RL fine-tuning of OpenVLA LoRA adapters on RL4VLA. It includes ManiSkill, SimplerEnv, openvla, real2sim, sim2real, and experiment launch scripts under scripts. The structure is typical for modern VLA experiments: model code and SFT live around openvla, environments and RL loops live around ManiSkill/SimplerEnv, and paper experiments are wrapped by shell scripts.

Task setup: RL for VLA manipulation
The paper models control as a partially observable Markov decision process. At each timestep, the policy receives:
- the current RGB observation
I_t; - a natural-language instruction
l, such as "put the carrot on the plate"; - no direct access to the full simulator state.
The policy predicts tokenized actions u_t. The environment wrapper converts those tokens into executable robot commands a_t. Episodes have a fixed horizon of T = 80, and the reward is a sparse difference-based signal derived from task progress. This is not a dense shaping reward where every tiny motion is manually scored.
The backbone is OpenVLA-warmup. The base model is frozen, and training updates a rank-32 LoRA adapter. For RL, the system uses the RL4VLA value-head modification so PPO can estimate advantages. The main task is PutOnPlateInScene25Main-v3: a simulated robot must place an object on a plate.
For beginners, the most important detail is that this is not full fine-tuning of a 7B model. It is parameter-efficient adaptation. That matters because full fine-tuning would change the compute cost, stability properties, and catastrophic forgetting behavior.
Main idea: offline supervision as a PPO prior
Standard PPO updates a policy using the likelihood ratio between the current policy and the old rollout policy:
r_t(theta) = pi_theta(a_t | o_t) / pi_theta_old(a_t | o_t)
The clipped PPO objective prevents updates from becoming too large. The problem is that robot manipulation rewards are often sparse. Early PPO can be almost blind: the policy has not learned useful action structure yet, successful episodes are rare, and exploration is inefficient.
Offline Supervision VLA-RL adds an auxiliary loss to PPO:
L(theta) = L_PPO(theta) + beta * L_aux(theta)
The auxiliary loss can be defined in two ways:
L_RefKL: keep the current policy close to a frozen SFT reference policy.L_DataBC: keep the current policy close to offline demonstration actions through behavior cloning.
The coefficient beta is scheduled. The paper holds beta = beta_0 for the first 100k steps, linearly decays it to zero over the next 200k steps, and then continues as pure PPO after 300k steps. Intuitively, the demonstrations guide the early phase, then the constraint fades so RL can optimize the final policy.

How RefKL works
RefKL uses an SFT policy as a frozen reference model. During RL, the current policy is penalized if its action distribution moves too far away from that reference:
L_RefKL = average KL(pi_ref(. | o) || pi_theta(. | o))
The full objective becomes PPO plus the weighted KL term. If the current policy starts drifting away from useful demonstration behavior, the KL penalty pulls it back. But the penalty is soft, so PPO can still move the policy when online reward shows that a different behavior is better.
The advantage of RefKL is that it uses the full action distribution of the reference policy. For a large VLA, that distribution can contain richer information than a single demonstrated action. A demonstration may show one valid motion, but the reference policy may encode a smoother set of nearby preferences.
The paper finds RefKL to be the most consistent guided variant. It stays above standard PPO across training curves and is generally more reliable than DataBC. This makes sense: KL-to-reference is a smooth regularizer, while BC on a fixed offline dataset can over-anchor the policy to states seen in the dataset.
How DataBC differs
DataBC does not need a reference-policy forward pass for every offline supervision signal. It directly adds a behavior cloning loss:
L_DataBC = - E log pi_theta(a_star | o)
Here (o, a_star) comes from the offline demonstration dataset. This is simple and easy to implement if your training stack already has an offline dataloader. It is also close to traditional supervised fine-tuning.
The downside is that the dataset only contains states visited by the demonstrator or motion planner. During online RL, the policy will visit new states, especially after mistakes. In those states, a fixed dataset BC term may be less informative and may sometimes pull the policy back toward the old distribution.
DataBC still beats PPO at the same 1M-step budget, but it is not as consistently strong as RefKL. If you need a quick baseline and have limited memory, DataBC is a good starting point. If you have a strong SFT checkpoint and enough GPU memory to evaluate it, RefKL is the stronger default.
Installation
The commands below follow the official repository README. The environment assumes CUDA 12.1 and PyTorch 2.2. You will need a Linux machine with a sufficiently large NVIDIA GPU for OpenVLA 7B LoRA work, or a multi-GPU setup if you plan to reproduce multiple seeds.
git clone https://github.com/alstar8/offline-supervision-vla-rl.git
cd offline-supervision-vla-rl
# Fast path: create the rlvla-guided environment
bash ./install_env.sh --flash-attn-wheel /path/to/flash_attn.whl
conda activate rlvla-guided
Manual installation:
conda create -n rlvla-guided python=3.10 -y
conda activate rlvla-guided
pip install torch==2.2.0 torchvision==0.17.0 torchaudio==2.2.0 \
--index-url https://download.pytorch.org/whl/cu121
cd openvla && pip install -e . && cd ..
pip install -U tyro datasets==3.3.2
cd ManiSkill && pip install -e . && cd ..
cd real2sim && pip install -e . && cd ..
cd SimplerEnv && pip install -e . && cd ..
The RLDS dataset builder uses a separate environment:
cd openvla/rlds_dataset_builder
conda env create -f environment_ubuntu.yml
conda activate rlds_env
For beginners, the likely failure points are CUDA/PyTorch compatibility, flash-attn, and editable installs across submodules. Debug layer by layer: first import torch, then mani_skill, then SimplerEnv, and only then launch training.
Preparing demonstration data
The experiments use PutOnPlateInScene25Main-v3. The repository collects demonstrations with the ManiSkill motion planner:
conda activate rlvla-guided
cd ManiSkill
python -m mani_skill.examples.motionplanning.widowx.collect_simpler \
-e "PutOnPlateInScene25Main-v3" \
--save_video \
--save_data \
--num_procs 16 \
--num_traj 16400 \
--seed=100
Then build the SFT RLDS dataset:
conda activate rlds_env
cd openvla/rlds_dataset_builder/sft_dataset
tfds build --overwrite
mkdir -p ../../../datasets
mv -T ~/tensorflow_datasets/example_dataset ../../../datasets/sft
The paper uses an early-stopped SFT reference checkpoint trained on 2k demonstrations for 7.5k steps. That detail is useful: the reference policy does not need to be a huge, fully converged SFT model. A reasonably good supervised checkpoint can already provide a valuable RL prior.
Training: SFT, PPO, RefKL, DataBC
The main scripts are:
| Method | Script |
|---|---|
| SFT reference | scripts/train_sft.sh |
| Standard PPO | scripts/train_ppo.sh |
| PPO SFT-init | scripts/train_ppo_sft_init.sh |
| RefKL | scripts/train_refkl.sh |
| DataBC | scripts/train_databc.sh |
Train the SFT reference:
bash scripts/train_sft.sh
Run the PPO baseline:
bash scripts/train_ppo.sh --seed 0
Run PPO initialized from the SFT LoRA:
bash scripts/train_ppo_sft_init.sh --seed 0
Run the two guided methods:
# RefKL: PPO + KL penalty to frozen SFT reference
bash scripts/train_refkl.sh --seed 0
# DataBC: PPO + offline behavior cloning auxiliary loss
bash scripts/train_databc.sh --seed 0
To cap training at 1M environment steps:
bash scripts/train_refkl.sh --seed 0 --steps_max=1000000
According to the repository mapping, RefKL corresponds to train_ms3_ppo_bc_teacher.py with --bc_to_ref_enabled, DataBC corresponds to train_ms3_ppo_sft.py, and the beta curriculum uses flags such as --bc_to_ref_coef, --bc_to_ref_hold_steps=100000, and --bc_to_ref_decay_steps=300000.
Inference and evaluation
In this paper, "inference" mainly means evaluating a trained checkpoint inside RL4VLA/SimplerEnv. It is not a direct real-robot deployment recipe. Set the checkpoint path and run:
CKPT_PATH="gen-robot/openvla-7b-rlvla-warmup" \
UNNORM_KEY="bridge_orig" \
VLA_LOAD_PATH="../SimplerEnv/wandb/<run>/glob/steps_XXXXXX" \
bash scripts/eval_policy.sh
Then aggregate success rates:
cd SimplerEnv/scripts
python calc_statistics.py
The evaluation protocol is:
- 64 in-distribution episodes per checkpoint.
- 960 out-of-distribution episodes in total.
- 15 OOD environments, 64 episodes each.
- OOD settings grouped into action, language, and vision shifts.
- Evaluation seeds
{0, 1, 2}.

Key results
The main comparison:
| Method | IND | OOD Act | OOD Lang | OOD Vis | OOD Avg |
|---|---|---|---|---|---|
| SFT | 0.82 | 0.46 | 0.60 | 0.74 | 0.62 |
| PPO SFT-init (1M) | 0.87 | 0.51 | 0.69 | 0.78 | 0.69 |
| PPO (2M) | 0.92 | 0.82 | 0.75 | 0.76 | 0.77 |
| PPO RefKL (1M) | 0.93 | 0.79 | 0.76 | 0.76 | 0.77 |
| PPO DataBC (1M) | 0.91 | 0.74 | 0.73 | 0.74 | 0.74 |
The headline result is simple: PPO RefKL at 1M steps reaches an OOD average of 0.77, matching PPO at 2M steps. Its IND score is also slightly higher, 0.93 versus 0.92. DataBC is also strong: 0.74 OOD average at 1M steps.
At the same 1M budget:
| Method | IND | OOD Avg |
|---|---|---|
| PPO (1M) | 0.75 | 0.64 |
| PPO SFT-init (1M) | 0.87 | 0.69 |
| PPO RefKL (1M) | 0.93 | 0.77 |
| PPO DataBC (1M) | 0.91 | 0.74 |
RefKL improves OOD average by +0.13 over PPO at 1M steps. DataBC improves it by +0.10. For robotics, those gains are substantial because every point of success rate usually costs many more rollouts.

Why not just initialize from SFT?
The obvious baseline is reasonable: train SFT, initialize PPO from the SFT LoRA, and continue with online RL. The paper shows that this is better than plain PPO at 1M steps, but it is still weaker than RefKL and DataBC.
One interpretation is that after the LoRA adapter has fit the supervised distribution strongly, online RL has a harder time changing the action behavior, especially under action OOD shifts. The policy starts from a better point, but its later adaptation can be slow.
RefKL is different. It does not only start from SFT. It injects supervised knowledge into the objective while keeping PPO active. Because beta decays to zero, the policy is guided early but still becomes a pure RL policy later. For deployment teams, the lesson is important: an SFT checkpoint is not only an initialization; it can also be a teacher during RL.
When should you use this?
Use Offline Supervision VLA-RL when:
- you already have demonstration data from teleoperation or motion planning;
- your task has sparse rewards, such as pick-place, insertion, drawer opening, or tool use;
- PPO or GRPO is too slow during the first 20-30% of training;
- you adapt a large VLA such as OpenVLA with LoRA or another PEFT method;
- you care about OOD generalization, not only one in-distribution scene.
Be careful when:
- demonstrations are low-quality or mismatched with the online task;
- the reward encourages shortcuts;
- you do not have enough memory to run a frozen reference policy;
- your environment reset or evaluation protocol is very different from RL4VLA.
If memory is tight, start with DataBC. If you can afford the reference forward pass, use RefKL as the stronger default.
Reproduction checklist
A practical checklist:
- Clone the repository and create
rlvla-guided. - Create the separate
rlds_envfor dataset building. - Collect demonstrations for
PutOnPlateInScene25Main-v3. - Build the RLDS SFT dataset.
- Train the SFT reference or use the checkpoint path expected by the scripts.
- Run PPO for 1M steps to establish your local baseline.
- Run RefKL for 1M steps with beta curriculum.
- Run DataBC for 1M steps as a cheaper guided baseline.
- Evaluate IND and all 15 OOD environments with the same seeds.
- Compare OOD average, not just IND success.
During early debugging, you can reduce seeds to save time. For a report or internal benchmark, keep the evaluation protocol close to the paper. If you only look at one in-distribution scene, you may miss the entire point of the method.
Connection to LeRobot and real robots
This method is not a drop-in LeRobot recipe yet, but the idea transfers cleanly. In a LeRobot-style workflow, you often have:
- an offline dataset collected by teleoperation;
- an SFT or behavior cloning policy;
- a simulator or real-robot online learning loop;
- a reward from success detection, a scripted metric, human intervention, or a learned evaluator.
Instead of throwing away the dataset after SFT, keep it active during online RL. If you have a frozen SFT policy, add KL-to-reference like RefKL. If not, sample batches from the dataset and add a BC loss like DataBC. Our HIL-SERL with LeRobot guide is a good companion for the online RL and human-intervention side; Offline Supervision VLA-RL adds the idea that demonstration data should stay alive inside the objective.
Conclusion
Offline Supervision VLA-RL matters because it targets the core pain point of robotics: online RL is powerful but expensive. The paper shows that offline demonstrations should not be used only once for SFT. When offline supervision is added to PPO through RefKL or DataBC, the policy learns faster, keeps OOD generalization, and RefKL can match the OOD average of 2M-step PPO with only 1M environment steps.
For beginners, the big lesson is to treat a demonstration dataset as a continuing training prior, not a file you use once and discard. For VLA manipulation teams, this is a practical route to reduce compute and interaction cost before moving toward real robot training.



