A robot following “put the bowl in the drawer” should not spend the same language-backbone computation on every background patch as on the bowl, gripper, and drawer edge. However, deleting the wrong patch can change the predicted action even when the robot still recognizes the object. VLA-ACL learns visual token selection from its effect on actions, giving the pruning policy a supervision signal directly connected to control.
This tutorial follows the original VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models paper, by Owen Du and colleagues, first submitted on October 6, 2026, and the official du-owen/VLA-ACL repository. We will explain the architecture, install the implementation, evaluate released checkpoints, train a patch scorer, and interpret LIBERO results. Benchmark numbers are reported by the authors; this article does not claim that vnrobo independently ran the GPU experiments.
You start with an already fine-tuned OpenVLA-OFT policy. Its vision encoder, language backbone, proprioception projector, and action head remain frozen. Only a patch scorer learns new weights. For another approach to reducing the visual information reaching an action pathway, see our explanation of Bottleneck Tokens in Pelican-VLA.
1. Understand the inputs before installing anything
A Vision-Language-Action model, or VLA, maps camera observations, a language instruction, and robot state to actions. LIBERO actions have seven components: three translation components, three orientation components, and a gripper command. Proprioception describes the robot's own state, such as end-effector pose or gripper state. It is a separate input rather than another camera image.
A visual token is a vector representing an image patch after the vision encoder. With multiple cameras, these vectors can dominate the input sequence processed by the language backbone. Fewer input tokens can reduce attention work and other computations performed for each token. VLA-ACL still encodes the original images before choosing patches, so this technique does not remove all vision-encoder computation or simply lower image resolution.
LIBERO is a simulated manipulation benchmark with Spatial, Object, Goal, and Long suites. LIBERO-Long is also called LIBERO-10. Its evaluation name is libero_10, whereas the training dataset used here is libero_10_no_noops. Keep these names separate: an evaluation suite selects environments, while a dataset name selects demonstrations for training.
For your first experiment, the intended outcome is modest and concrete: load a known policy, run a short baseline rollout, enable the matching scorer, and confirm that both complete without configuration errors. Full benchmark reproduction comes after this pipeline works.
2. The paper's idea: select patches that preserve actions
Imagine that a full-context policy already produces useful actions. Run it with all visual tokens to obtain a teacher prediction. Then run the same frozen policy with a differentiably modified visual context. Train a scorer so that the modified-context actions stay close to the teacher and to actions recorded in demonstrations.
ACL stands for Action Consistency Learning. The consistency target is the action output, rather than an object category, segmentation mask, or internal attention map. This means the scorer can learn from ordinary robot demonstrations without requiring hand-labeled important patches.
Consider a small patch containing the contact region between a gripper and a handle. It may have less semantic prominence than a large colorful background region, yet removing it could change the action. Training against downstream action predictions allows the scorer to discover this distinction. Attention scores can be useful heuristics, but they do not directly measure whether deleting a token changes control output.

The visualization compares pruning and caching over a trajectory. VLA-ACL applies its budget at every policy query, including the first query. It does not depend on a temporal controller periodically restoring full context. That makes the inference path predictable, but a small budget must remain adequate throughout the entire task.
3. Architecture: score jointly, select separately for each view
In the LIBERO setup, two camera views each produce 256 visual tokens, giving 512 image tokens before pruning. The scorer is a five-layer bidirectional Transformer with hidden dimension 1024 and 16 attention heads. A linear head produces one logit per patch. The patch scorer implementation takes vision embeddings before their projection into the LLM embedding space.
The scorer processes concatenated embeddings from both views. Token selection then uses an independent Top-K budget for each camera. At K=64, it retains 64 patches from the external view and 64 from the wrist view: 128 visual tokens in total. K does not describe the combined budget for all cameras.
| Configuration | Tokens per view | Image tokens across two views | Image-token pruning rate |
|---|---|---|---|
| Full context | 256 | 512 | 0% |
| K=64 | 64 | 128 | 75% |
| K=32 | 32 | 64 | 87.5% |
These counts exclude instruction, proprioception, and action tokens. Removing 87.5% of image tokens therefore does not mean eliminating 87.5% of the complete inference workload.
Position handling is also important. Retained tokens preserve their original RoPE position indices. The implementation selects by score and sorts the chosen indices back into their original sequence order. If you write your own token-gathering code, simply renumbering the shortened sequence can change how the frozen backbone interprets positional relationships.
Two images -> frozen vision encoder -> vision embeddings
|
trainable patch scorer
|
Top-K within each view
|
Instruction + proprio -> frozen language backbone -> action chunk
The scorer is an additional module, not a replacement for the base policy. The action head and normalization statistics still belong to the policy checkpoint and must be loaded correctly.
4. Soft selection during training, hard pruning during inference
Hard Top-K makes discrete choices. Small changes in a patch logit usually leave the selected set unchanged, so selection has zero gradients almost everywhere. Directly training a scorer through this operation would make it difficult for action loss to identify how scores should change.
VLA-ACL uses a soft Top-K relaxation during training. Each patch receives a selection score between zero and one, with scores summing to K within each view. An adaptive threshold enforces this budget, and a Laplace cumulative distribution function maps scaled logit distances into scores. The softness parameter alpha controls how closely this resembles hard selection.
The soft gate combines the original patch embedding with an embedding obtained by encoding a black image through the same preprocessing and vision encoder:
X_soft[i] = s[i] * X[i] + sqrt(1 - s[i]^2) * black_embedding[i]
loss = 0.5 * mean_absolute_error(action_soft, action_teacher)
+ 0.5 * mean_absolute_error(action_soft, action_demo)
This block explains the algorithm; it is not a standalone training program. A score close to one preserves the original patch. A score close to zero substitutes the corresponding black-image embedding. Training keeps the token positions available for differentiable supervision; inference actually removes unselected tokens before they enter the language backbone.

Why use black-image embeddings? Multiplying embeddings by a score until they become zero vectors creates a different surrogate. The paper compares gating variants and finds that its black-image gate better narrows the soft-to-hard action gap in the tested configuration. The implementation detaches the black-image weight when calculating gradients because the square-root derivative becomes large near a score of one; gradients still reach the scorer through the s * X branch.
A frozen backbone must still support gradients with respect to the soft-gated inputs. You can run the teacher under torch.no_grad(), but placing the entire soft-pruned forward under it would prevent scorer learning. The implementation freezes backbone parameters while allowing input gradients and uses gradient checkpointing for that forward path. This explains why training a small scorer can still require substantial activation memory.
5. Install the VLA-ACL fork
Prepare Linux, Conda, a compatible NVIDIA GPU and CUDA environment, and enough disk space for the policy, datasets, and download caches. The experiments use one A100. The repository does not establish that the default batch size of eight fits every smaller GPU; a lightweight scorer does not eliminate the memory occupied by a 7B policy.
Run these commands on your training machine, starting outside the blog repository:
conda create -n vla-acl python=3.10 -y
conda activate vla-acl
git clone https://github.com/du-owen/VLA-ACL.git
cd VLA-ACL/src/openvla-oft
# CUDA 12.1 is an example; choose wheels compatible with your driver.
pip install torch==2.2.0 torchvision==0.17.0 torchaudio==2.2.0 \
--index-url https://download.pytorch.org/whl/cu121
pip install -e .
pip install packaging ninja
pip install "flash-attn==2.5.5" --no-build-isolation
git clone https://github.com/Lifelong-Robot-Learning/LIBERO.git
pip install -e LIBERO
pip install -r experiments/robot/libero/libero_requirements.txt
The repository's SETUP.md inherits upstream OpenVLA-OFT instructions and still mentions cloning the upstream project. For this tutorial, install the editable package from VLA-ACL's src/openvla-oft directory, because it contains the scorer and pruning-aware modeling code. Its pyproject.toml pins PyTorch 2.2.0 and uses a custom Transformers fork. Replacing that fork with an arbitrary newer release changes the environment being reproduced.
Check basic imports before downloading large files:
python -c "import torch; print(torch.__version__, torch.cuda.is_available())"
python -c "import transformers; print(transformers.__version__)"
python -c "from prismatic.models.patch_scorer import PatchScorer; print('scorer import OK')"
If CUDA is unavailable, fix the driver and PyTorch installation first. FlashAttention compilation problems require checking toolkit, compiler, and dependency compatibility. Random package upgrades can make imports work while breaking the intended model behavior, so keep a record of the versions you actually use.
6. Download the complete policy and matching scorer
Using the Python Hugging Face API avoids depending on a CLI command name that differs across package versions. From src/openvla-oft, download the LIBERO-10 policy snapshot and its K=64 scorer:
python - <<'PY'
from huggingface_hub import snapshot_download, hf_hub_download
snapshot_download(
repo_id="moojink/openvla-7b-oft-finetuned-libero-10",
local_dir="checkpoints/openvla-7b-oft-finetuned-libero-10",
)
hf_hub_download(
repo_id="owendudu/VLA-ACL",
filename="scorer-libero-10-k64.pt",
local_dir="checkpoints/vla-acl",
)
PY
The released scorers cover Spatial, Object, Goal, and 10 at K=32 and K=64. Each scorer state dictionary is approximately 130 MB. A scorer file is not a complete VLA policy: you still need the base checkpoint, action head, proprioception projector, and dataset statistics. Downloading the full policy snapshot is safer than selecting only its main weight shards.
For evaluation alone, skip the RLDS training dataset. To train a scorer, download it separately:
python - <<'PY'
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="openvla/modified_libero_rlds",
repo_type="dataset",
local_dir="modified_libero_rlds",
)
PY
LIBERO.md describes approximately 10 GB across the four suites. Check that the root contains libero_10_no_noops. The suffix means demonstrations with near-zero no-op actions were filtered. Pass the parent directory as --data_root_dir, not the suite directory itself.
7. Run a baseline and a short K=64 evaluation
First run a small evaluation to expose checkpoint, import, and rendering issues. On a headless server, MUJOCO_GL=egl may be appropriate when the GPU driver and EGL setup support it. Treat this as a rendering configuration choice rather than a universal fix.
python experiments/robot/libero/run_libero_eval.py \
--pretrained_checkpoint checkpoints/openvla-7b-oft-finetuned-libero-10 \
--task_suite_name libero_10 \
--center_crop True \
--num_trials_per_task 2 \
--seed 67 \
--patch_scorer.enabled False
Next enable pruning while keeping the same suite, seed, and image preprocessing:
python experiments/robot/libero/run_libero_eval.py \
--pretrained_checkpoint checkpoints/openvla-7b-oft-finetuned-libero-10 \
--task_suite_name libero_10 \
--center_crop True \
--num_trials_per_task 2 \
--seed 67 \
--patch_scorer.enabled True \
--patch_scorer.checkpoint checkpoints/vla-acl/scorer-libero-10-k64.pt \
--patch_scorer.prune_topk 64
The local checkpoint path is important. VLA-ACL's README asks for a locally stored policy so the loader can synchronize modeling code with this fork. It updates AutoMap configuration and checks the local modeling files. Passing only the Hub model ID during evaluation can load older modeling logic without patch_scorer support in predict_action, even when your scorer weights are correct.
The scorer must match the policy suite, and prune_topk must match the K encoded in the scorer filename. Do not pair a Spatial scorer with the LIBERO-10 policy or assume that evaluating a K=64-trained scorer at K=32 reproduces the released K=32 configuration.
Two episodes per task are a smoke check, not a reliable accuracy estimate. Once both paths run, use 50 trials per task and repeat across three seeds, matching seeds between baseline and pruning evaluations. The paper averages 1,500 episodes per suite: ten tasks times 50 trials times three seeds. Record your actual seed list; the example seed 67 alone does not reproduce that protocol.
8. Train your own patch scorer
The following command uses the paper's main training settings on LIBERO-10:
python vla-scripts/train_scorer.py \
--vla_path checkpoints/openvla-7b-oft-finetuned-libero-10 \
--data_root_dir modified_libero_rlds \
--dataset_name libero_10_no_noops \
--soft_top_k 64 \
--batch_size 8 \
--learning_rate 0.0001 \
--max_steps 5000 \
--teacher_weight 0.5 \
--gt_weight 0.5 \
--use_wandb False
Disabling WandB avoids the script's placeholder entity and project defaults. Enable it with your own settings if you want dashboard metrics. The training script uses AdamW and a cosine learning-rate schedule. Selection softness alpha follows a cosine schedule from 2.0 to 0.1 over 5,000 steps.
If training exceeds available VRAM, reduce batch size and document the change. --grad_accumulation_steps lets you adjust effective batch size, but cannot solve a GPU that lacks enough memory for the model and a single sample. Memory figures for upstream policy fine-tuning should not be presented as guaranteed scorer-training requirements because the computation graphs differ.
With the shown defaults, the final scorer is saved under:
runs/vla-acl+libero_10_no_noops+k64+teacher-0.5+gt-0.5+b8+lr-0.0001/
patch_scorer_head--5000_checkpoint.pt
Changing batch size, accumulation, learning rate, or a run note changes the directory name. Use the exact saved path printed by the script. Evaluate it by replacing --patch_scorer.checkpoint in the earlier command, while keeping --patch_scorer.prune_topk 64. Train or download the corresponding K=32 scorer before switching to that configuration.
When logging is enabled, inspect total action loss, teacher loss, demonstration loss, and hard Top-K metrics. Low soft-gate loss alongside high hard-pruning loss suggests that training's surrogate is not closely matching inference. Check the alpha schedule, checkpoint architecture, and preprocessing before deciding that the entire VLA needs retraining. For a different learning intervention on the same benchmark family, see trajectory-level VLA training with TGRPO.
9. Interpret the results without confusing FLOPs and latency
The paper's Table I reports the following A100 results. Its TFLOPs and latency metrics measure the LLM, rather than the complete observation-to-actuation loop.
| Method | Spatial | Object | Goal | Long | Average | TFLOPs | LLM latency |
|---|---|---|---|---|---|---|---|
| OpenVLA-OFT | 97.8% | 97.6% | 97.6% | 94.2% | 96.8% | 4.013 | 63.49 ms |
| VLA-ACL K=64 | 98.7% | 98.4% | 96.5% | 94.2% | 97.0% | 1.412 | 46.82 ms |
| VLA-ACL K=32 | 98.2% | 98.2% | 96.6% | 88.7% | 95.4% | 0.991 | 42.30 ms |

At K=32, FLOPs decrease by approximately 75.3%, while latency improves by 63.49 / 42.30 = 1.50×. The difference illustrates why computational savings do not translate linearly into runtime savings. K=64 gives approximately 1.36× from the table, rounded to 1.4× in the paper.
The most important accuracy tradeoff is on Long. K=32 loses 5.5 percentage points relative to the baseline, whereas K=64 retains 94.2%. For a beginner, K=64 is therefore a sensible first configuration: verify task-level performance, then experiment with the more aggressive budget. Its average improvement of 0.2 points does not by itself establish a statistically significant gain on every setup.
Baseline success rates in the paper are taken from their corresponding publications. Do not describe this table as evidence that the authors independently retrained and reran every comparison method. For your own study, measure policy latency, end-to-end latency, VRAM, task success, and failure categories on the same hardware and preprocessing.
LIBERO predicts chunks of eight actions in this setup. One policy forward is not equivalent to one actuator control tick. Converting an LLM latency into a robot control frequency without accounting for perception, action execution, and chunking would be misleading.
10. What the real-world and π0.5 experiments establish
The paper also evaluates four tasks on an ALOHA-style AgileX PiPER-X platform. At K=64, average success is 72.5% versus 73.75% for OpenVLA-OFT, while measured LLM latency falls from 91.5 to 62.1 ms. The base policies are fine-tuned using 50–100 expert demonstrations per task. These results do not mean that a released LIBERO scorer transfers directly to a physical robot without task-specific preparation.
The appendix extends action consistency learning to π0.5. Because that model uses flow matching, supervision compares predicted velocities at the same noise level, with the flow-matching target providing the auxiliary signal. Its reported 2.2× improvement measures LLM pre-fill latency. The public README's runnable scorer workflow focuses on OpenVLA-OFT, and train_scorer.py explicitly supports continuous L1 regression heads. Do not assume that the same command trains the appendix's π0.5 variant.
Before deployment, verify camera order, normalization, and the action contract for the target robot. Our VLA PEFT and deployment guide places these checks in a broader policy workflow. Preserving a teacher's action predictions cannot automatically correct every weakness of that teacher or guarantee behavior on unfamiliar observations.
11. Troubleshoot the first complete experiment
| Symptom | First thing to check |
|---|---|
| Patch scorer import fails | Does the active environment contain the editable VLA-ACL fork? |
predict_action lacks pruning support |
Are you evaluating a local policy snapshot with synchronized modeling code? |
| Scorer state dictionary has shape mismatches | Do hidden dimension 1024, five layers, and 16 heads match training? |
| Dataset or statistics key is missing | Do the dataset root, _no_noops name, and policy suite agree? |
| Training stops during logging initialization | Is WandB disabled or configured with your own entity and project? |
| Simulator cannot render observations | Does the selected MuJoCo graphics backend work with the GPU driver? |
| K=32 is faster but long tasks fail | Compare K=64 with matched seeds and inspect rollout videos for lost context |
A useful experiment record includes policy identity, scorer filename, K, package versions, seed list, image preprocessing, and both task-level success and latency. Keep these details beside your results. They let you distinguish a pruning tradeoff from a checkpoint mismatch or an environment change.
The practical sequence is to establish a full-context baseline, validate the released K=64 scorer, train a new scorer only when needed, and then expand evaluation to enough trials to assess accuracy. Inspecting individual failures is especially useful on Long: a small change in action prediction can compound over several subtasks even when the average metric looks acceptable.



