Quick Summary
UMR/WEPVLA is worth studying because it attacks one of the hardest practical questions in robot manipulation: can we train useful manipulation policies from human demonstrations instead of collecting endless robot trajectories? The paper UMR: Universal Manipulation Representation, the project page at umr-wepvla.github.io, and the open-source repo LiuSong-Scrat/UMR answer this with a geometric representation rather than another generic VLA wrapper.
UMR represents the same manipulation motion through two linked views. World Flow describes task-relevant object motion in the world frame. Ego Trajectory describes the end-effector motion relative to its current pose, which is closer to the command a robot controller can execute. The two views are coupled by an SE(3) conjugation, so they are not two unrelated labels. They are two coordinate descriptions of the same physical motion.
The authors instantiate this representation as World-Ego Point VLA (WEPVLA), a compact point-cloud VLA with about 0.5B parameters, a dual-stream Point Action Adapter, and a shared Point Action Expert. In the reported real-world experiments, a single policy trained with about 10 minutes of human demonstrations per task, and no robot demonstrations, reaches 91.7% average success across six deployment settings. The HumanEgo baseline reaches 60.8% under the same setting. In simulation, the project page reports 85.7% on the 10-task RLBench benchmark and 97.5% across LIBERO's four suites.
If you already know OpenVLA, SmolVLA training in LeRobot, or the practical UMI/VLA workflow, think of UMR as a geometric action language that sits between human videos, robot data, and executable action chunks.
The Problem: Human Demonstrations Are Rich but Not Robot Actions
A person stacking cubes, folding a towel, sweeping trash, or placing a mug into a rack provides a lot of useful information. The demonstration shows which object matters, where contact happens, how the object should move, and when recovery is needed. But the demonstration does not directly provide robot actions. Human hands do not share a robot arm's kinematics. A first-person camera does not share the exact deployment extrinsics of a robot camera. A human wrist path is not a joint command or end-effector command for a robot.
Most human-to-robot approaches tend to fall into one of three buckets:
| Approach | Core idea | Common weakness |
|---|---|---|
| Direct retargeting | Estimate human motion and map it to the robot | Sensitive to embodiment, grasp, and calibration |
| World-only motion | Learn object or point flow independent of embodiment | Needs a separate module to turn world motion into robot actions |
| Robot-only VLA | Train directly on robot demonstrations | Expensive to scale across tasks and embodiments |
UMR sits between these options. It does not discard world motion, because object motion transfers better than joint commands. It also does not leave execution to a detached post-processor, because the policy still learns an executable Ego Trajectory. The World-Ego pair lets the model learn both "where the task wants the object to go" and "how the current gripper should move from here."

The Paper Idea: One Motion, Two Views
Consider a simple laptop-closing task. In the world frame, the important motion is the laptop lid rotating around its hinge until it is closed. That is World Flow: a task-level physical effect. From the robot's current end-effector frame, the action is different: approach the lid edge, move downward along a feasible arc, keep contact, and stop at the right time. That is Ego Trajectory.
UMR says these are not independent targets. They are two coordinate views of the same physical motion. If the policy knows the transform between the world frame and the current end-effector frame, it can connect the two with SE(3) conjugation. This is the central trick. It lets task intent learned from one demonstrator be represented as executable motion for another embodiment.
For a beginner, the representation can be remembered like this:
World Flow:
"How should the task-relevant object move in the world?"
Ego Trajectory:
"How should this robot's end-effector move from its current pose?"
UMR:
Use SE(3) geometry to make those two answers describe the same motion.
This is also different from unconstrained dense point flow. UMR uses a compact SE(3) trajectory. Each action chunk contains translation, a continuous 6D rotation representation, and gripper state. That structure matters because it reduces ambiguity and makes the target easier for a smaller policy to learn.
WEPVLA Architecture
WEPVLA turns UMR into an end-to-end point-cloud VLA. The paper and project page describe three main pieces.
Motion-Aware Segmentation filters the point cloud toward interaction-relevant foreground. Instead of passing the entire scene to an action decoder and hoping attention discovers the correct object, the model reasons about points that move with the interaction versus points that remain static in the world. For manipulation, a few centimeters around the contact region often matter more than the background.
Dual-stream Point Action Adapter has a World stream and an Ego stream. The World stream represents observations and actions in the world frame. The Ego stream re-expresses them in the current end-effector frame. Both streams use a Coarse-to-Fine Point Encoder to produce fine foreground tokens and coarse background tokens. Action tokens are then fused with point features through bidirectional self-attention.
Shared Point Action Expert receives action tokens, scene tokens, and frozen VLM image/language tokens, then predicts both action streams. During training, the model learns World Flow and Ego Trajectory together. During inference, the two streams are jointly denoised, but only the Ego Trajectory is sent to the robot. This is a practical design: World Flow provides a transferable task prior, while Ego Trajectory provides the command interface.

Data-Efficient Strategy: More Variation Without Breaking Contact
The most practical part of the paper is the Data-Efficient Strategy (DES). With only about 10 minutes of human demonstrations per task, the dataset can be narrow. Objects appear in familiar positions, approach directions repeat, and recovery behavior is limited. Simple 2D image augmentation may create visual variation, but it does not guarantee that the contact geometry still makes sense.
DES augments data at the point-cloud level and respects the task stage. It identifies the interaction-relevant object and perturbs object configurations in 3D while preserving demonstrated contact geometry. For mug racking, the mug and rack can move to new tabletop positions, but the approach and relative geometry must remain feasible. For trash sweeping, the trash and dustpan can be repositioned, but the sweep-to-container relation must survive.
The project page reports that 100 human demonstrations can be expanded with 500 additional DES episodes. The important part is not cosmetic data multiplication. The useful part is exposing the policy to object configurations it did not see in the original human demos, while keeping the manipulation geometry valid.
Installing the Open-Source Repo
The current repo expects Python 3.10, an editable LeRobot-style install, RLBench/LIBERO benchmark code, and SmolVLM2/SmolVLA model resources. Plan for a Linux machine with an NVIDIA GPU, matching CUDA/PyTorch versions, enough storage for datasets, and CoppeliaSim if you run RLBench.
The basic setup from the README is:
sudo apt-get update
sudo apt-get install -y git git-lfs tmux xvfb xauth libegl1 libgl1-mesa-glx
git lfs install
cd /path/to/lerobot
conda create -n wepvla python=3.10 -y
conda activate wepvla
pip install -e ".[smolvla,libero]"
For the bundled RLBench implementation:
pip install -e benchmarks/RLBench
python -c "import pyrep, rlbench; print('RLBench import ok')"
Before starting a training run, verify that PyTorch sees the GPU:
python -c "import torch; print(torch.__version__, torch.cuda.is_available(), torch.cuda.device_count())"
nvidia-smi
Download the model resources:
bash benchmarks/RLBench/scripts/download_vlm_models.sh
The expected layout is:
benchmarks/vlm_model/
├── SmolVLM2-500M-Video-Instruct/
└── smolvla_base/
For offline clusters, pass paths directly:
VLM_MODEL_NAME=/path/to/SmolVLM2-500M-Video-Instruct \
VLM_WEIGHTS_PATH=/path/to/smolvla_base \
PYTHON=python \
bash benchmarks/RLBench/scripts/collect_data.sh \
--dataset-root /path/to/rlbench_dataset
One small but expensive warning from the README: CUDA extensions must be rebuilt for the target machine's PyTorch and CUDA versions. Copying local build artifacts between machines is a fast way to create confusing import errors.
Training on RLBench
RLBench requires a compatible CoppeliaSim installation. You must pass COPPELIASIM_ROOT, LD_LIBRARY_PATH, and Qt-related variables to the commands. On a headless server, start Xvfb:
Xvfb :99 -screen 0 1280x1024x24 -nolisten tcp >/tmp/rlbench-xvfb.log 2>&1 &
Collect data and build the cache:
PYTHON=python \
DATASET_ROOT=/path/to/rlbench_dataset \
COPPELIASIM_ROOT=/path/to/CoppeliaSim \
LD_LIBRARY_PATH="/path/to/CoppeliaSim:${LD_LIBRARY_PATH:-}" \
QT_QPA_PLATFORM=xcb \
QT_QPA_PLATFORM_PLUGIN_PATH=/path/to/CoppeliaSim \
QT_PLUGIN_PATH="" \
bash benchmarks/RLBench/scripts/collect_data.sh \
--dataset-root /path/to/rlbench_dataset
Train:
DATASET_ROOT=/path/to/rlbench_dataset \
OUTPUT_ROOT=/path/to/rlbench_output \
GPU_IDS=0 \
bash benchmarks/RLBench/scripts/train.sh
Evaluate the 10 tasks:
EVAL_POLICY_PATH=/path/to/checkpoint/pretrained_model \
EVAL_ROOT=/path/to/rlbench_eval \
EVAL_SAVE_VIDEO=0 \
EVAL_SAVE_ACTION_RECORDS=0 \
EVAL_SAVE_ACTION_CHUNKS=0 \
DISPLAY=:99 \
bash benchmarks/RLBench/scripts/evaluate.sh \
--tasks close_box close_fridge close_laptop_lid phone_on_base stack_wine \
sweep_to_dustpan take_frame_off_hanger \
take_umbrella_out_of_umbrella_stand toilet_seat_down water_plants \
--episodes 100
If you are new to RLBench, run one task and one episode first. Enable video and action logging while debugging. Only run the full evaluation after the simulator, display, dataset, and checkpoint path are all boringly correct.

Training and Inference on LIBERO
The LIBERO entry points live under benchmarks/song_real_libero/. The flow is similar: convert demonstrations, build the PointSeg cache, train, then evaluate.
Convert demonstrations:
PYTHON_BIN=/path/to/python \
DEMO_ROOT=/path/to/libero_demos \
DATASET_ROOT=/path/to/libero_dataset \
bash benchmarks/song_real_libero/prepare_dataset.sh
Build the cache:
PYTHON_BIN=/path/to/python \
DATASET_ROOT=/path/to/libero_dataset \
CACHE_ROOT=/path/to/libero_cache \
GPU_IDS=0 NPROC=1 \
bash benchmarks/song_real_libero/build_cache.sh
Train:
PYTHON_BIN=/path/to/python \
DATASET_ROOT=/path/to/libero_dataset \
CACHE_ROOT=/path/to/libero_cache \
BASE_POLICY=/path/to/base_policy/pretrained_model \
OUTPUT_ROOT=/path/to/libero_output \
GPU_IDS=0 \
bash benchmarks/song_real_libero/train.sh
Evaluate:
PYTHON_BIN=/path/to/python \
POLICY_PATH=/path/to/checkpoint/pretrained_model \
OUTPUT_DIR="benchmarks/song_real_libero/outputs/eval_$(date +%Y%m%d_%H%M%S)" \
CUDA_DEVICE=0 EPISODES=50 \
bash benchmarks/song_real_libero/evaluate.sh
At inference time, remember that this is not a plain image-to-action policy. The system builds a point-cloud observation, runs motion-aware segmentation, passes tokens through World and Ego streams, jointly denoises an action chunk, and executes the Ego action. For real deployment, log camera/depth health, point-cloud shapes, gripper state, predicted chunks, executed Ego trajectory, and task outcomes.

How to Read the Results
The headline numbers are strong, but they should be read in scope. On simulation, WEPVLA reports 85.7% mean success on the 10 RLBench tasks, compared with 82.3% for PointACT and 74.0% for HybridVLA in the project page table, while using a smaller 0.5B model. On LIBERO, WEPVLA reports 97.5%, compared with 96.0% for PointACT and 96.8% for OpenVLA-OFT.
The real-world comparison is more interesting. WEPVLA and HumanEgo are both trained with human-only data, about 10 minutes per task. WEPVLA reports 95.0% cross-environment success, 91.7% cross-embodiment success, 90.0% cross-setup success, and 91.7% average success across Tasks I-VI. HumanEgo reports 70.0%, 60.8%, 58.3%, and 60.8% under the same categories.
That does not mean a phone video plus 10 minutes of recording gives you a production robot. The reported system still depends on RGB-D sensing, point clouds, calibrated robot setups, task framing, and a controlled evaluation protocol. The useful lesson is narrower and stronger: if the action representation is geometric enough, human demonstrations become much more useful for robot policy learning.
When Should You Try UMR/WEPVLA?
UMR/WEPVLA is worth trying if your team has at least some of the following:
- RGB-D or multi-view data, or the discipline to set up calibration properly.
- Tabletop manipulation tasks with clear object motion: stacking, folding, sweeping, placing, racking, or closing.
- A desire to compare human demonstrations with robot teleoperation, instead of assuming robot-only data is the only path.
- Existing LeRobot, SmolVLA, or OpenVLA experiments that need better 3D grounding.
- Enough GPU budget to run real training and evaluation, not just a notebook demo.
Be slower if the task requires heavy tactile feedback, unstable deformable-object handling, force-sensitive closed-loop control, or mobile manipulation with poor depth sensing. WEPVLA improves the representation. It does not remove the need for calibration, controller safety, dataset QA, or failure analysis.
Beginner Checklist
- Read the paper sections on UMR, WEPVLA, and DES before changing the code.
- Run evaluation with an available checkpoint before training your own model.
- Verify dataset paths, model paths, CUDA, and simulator imports separately.
- For RLBench, start with one task and one episode, with video/action logs enabled.
- For LIBERO, write each evaluation to a fresh output directory.
- When training a custom task, log both World Flow predictions and executed Ego actions.
- If results are poor, inspect point clouds, frame transforms, and gripper state before changing model hyperparameters.
UMR/WEPVLA matters because it reframes the policy question. Should a manipulation model learn object motion or robot action? UMR's answer is: learn both, force them to agree through geometry, and execute the robot-local part. That is a compact idea, backed by code, benchmarks, and enough practical detail for small labs to start testing it seriously.



