VnRoboVnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
VnRobo logoVnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam
VnRoboVnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
  1. Home
  2. Blog
  3. UMR/WEPVLA: World-Ego Point VLA
wholebody-vlavlawepvlaumrhuman-demopoint-cloudmanipulation

UMR/WEPVLA: World-Ego Point VLA

A beginner guide to open-source UMR/WEPVLA: robot manipulation from 10 minutes of human demos using World Flow, Ego Trajectory, and point clouds.

Nguyễn Anh TuấnOctober 5, 202611 min readUpdated: Oct 9, 2026
UMR/WEPVLA: World-Ego Point VLA

Quick Summary

UMR/WEPVLA is worth studying because it attacks one of the hardest practical questions in robot manipulation: can we train useful manipulation policies from human demonstrations instead of collecting endless robot trajectories? The paper UMR: Universal Manipulation Representation, the project page at umr-wepvla.github.io, and the open-source repo LiuSong-Scrat/UMR answer this with a geometric representation rather than another generic VLA wrapper.

UMR represents the same manipulation motion through two linked views. World Flow describes task-relevant object motion in the world frame. Ego Trajectory describes the end-effector motion relative to its current pose, which is closer to the command a robot controller can execute. The two views are coupled by an SE(3) conjugation, so they are not two unrelated labels. They are two coordinate descriptions of the same physical motion.

The authors instantiate this representation as World-Ego Point VLA (WEPVLA), a compact point-cloud VLA with about 0.5B parameters, a dual-stream Point Action Adapter, and a shared Point Action Expert. In the reported real-world experiments, a single policy trained with about 10 minutes of human demonstrations per task, and no robot demonstrations, reaches 91.7% average success across six deployment settings. The HumanEgo baseline reaches 60.8% under the same setting. In simulation, the project page reports 85.7% on the 10-task RLBench benchmark and 97.5% across LIBERO's four suites.

Tool recommendations

VLA train/deploy stack

Train on cloud/workstation, then deploy optimized models to Jetson or the robot computer.

Cloud GPU for VLA / policy training Use for imitation learning, diffusion policies, RL, and robotics model fine-tuning. View cloud GPU → NVIDIA Jetson Orin NX / Orin Nano Edge deployment hardware for perception, logging, and optimized inference. View Jetson → Hugging Face / robotics dataset hosting Host datasets, checkpoints, and model cards for cleaner LeRobot/VLA workflows. View platform →
UMR/WEPVLA teaser: World-Ego action representation for human-to-robot manipulation — source: UMR project page

If you already know OpenVLA, SmolVLA training in LeRobot, or the practical UMI/VLA workflow, think of UMR as a geometric action language that sits between human videos, robot data, and executable action chunks.

The Problem: Human Demonstrations Are Rich but Not Robot Actions

A person stacking cubes, folding a towel, sweeping trash, or placing a mug into a rack provides a lot of useful information. The demonstration shows which object matters, where contact happens, how the object should move, and when recovery is needed. But the demonstration does not directly provide robot actions. Human hands do not share a robot arm's kinematics. A first-person camera does not share the exact deployment extrinsics of a robot camera. A human wrist path is not a joint command or end-effector command for a robot.

Most human-to-robot approaches tend to fall into one of three buckets:

Approach Core idea Common weakness
Direct retargeting Estimate human motion and map it to the robot Sensitive to embodiment, grasp, and calibration
World-only motion Learn object or point flow independent of embodiment Needs a separate module to turn world motion into robot actions
Robot-only VLA Train directly on robot demonstrations Expensive to scale across tasks and embodiments

UMR sits between these options. It does not discard world motion, because object motion transfers better than joint commands. It also does not leave execution to a detached post-processor, because the policy still learns an executable Ego Trajectory. The World-Ego pair lets the model learn both "where the task wants the object to go" and "how the current gripper should move from here."

WEPVLA pipeline from heterogeneous demonstrations to co-training and deployment — source: UMR project page
WEPVLA pipeline from heterogeneous demonstrations to co-training and deployment — source: UMR project page

The Paper Idea: One Motion, Two Views

Consider a simple laptop-closing task. In the world frame, the important motion is the laptop lid rotating around its hinge until it is closed. That is World Flow: a task-level physical effect. From the robot's current end-effector frame, the action is different: approach the lid edge, move downward along a feasible arc, keep contact, and stop at the right time. That is Ego Trajectory.

UMR says these are not independent targets. They are two coordinate views of the same physical motion. If the policy knows the transform between the world frame and the current end-effector frame, it can connect the two with SE(3) conjugation. This is the central trick. It lets task intent learned from one demonstrator be represented as executable motion for another embodiment.

For a beginner, the representation can be remembered like this:

text
World Flow:
  "How should the task-relevant object move in the world?"

Ego Trajectory:
  "How should this robot's end-effector move from its current pose?"

UMR:
  Use SE(3) geometry to make those two answers describe the same motion.

This is also different from unconstrained dense point flow. UMR uses a compact SE(3) trajectory. Each action chunk contains translation, a continuous 6D rotation representation, and gripper state. That structure matters because it reduces ambiguity and makes the target easier for a smaller policy to learn.

WEPVLA Architecture

WEPVLA turns UMR into an end-to-end point-cloud VLA. The paper and project page describe three main pieces.

Motion-Aware Segmentation filters the point cloud toward interaction-relevant foreground. Instead of passing the entire scene to an action decoder and hoping attention discovers the correct object, the model reasons about points that move with the interaction versus points that remain static in the world. For manipulation, a few centimeters around the contact region often matter more than the background.

Dual-stream Point Action Adapter has a World stream and an Ego stream. The World stream represents observations and actions in the world frame. The Ego stream re-expresses them in the current end-effector frame. Both streams use a Coarse-to-Fine Point Encoder to produce fine foreground tokens and coarse background tokens. Action tokens are then fused with point features through bidirectional self-attention.

Shared Point Action Expert receives action tokens, scene tokens, and frozen VLM image/language tokens, then predicts both action streams. During training, the model learns World Flow and Ego Trajectory together. During inference, the two streams are jointly denoised, but only the Ego Trajectory is sent to the robot. This is a practical design: World Flow provides a transferable task prior, while Ego Trajectory provides the command interface.

UMR real-world settings with robot embodiments, tabletop scenes, and gripper configurations — source: UMR project page
UMR real-world settings with robot embodiments, tabletop scenes, and gripper configurations — source: UMR project page

Data-Efficient Strategy: More Variation Without Breaking Contact

The most practical part of the paper is the Data-Efficient Strategy (DES). With only about 10 minutes of human demonstrations per task, the dataset can be narrow. Objects appear in familiar positions, approach directions repeat, and recovery behavior is limited. Simple 2D image augmentation may create visual variation, but it does not guarantee that the contact geometry still makes sense.

DES augments data at the point-cloud level and respects the task stage. It identifies the interaction-relevant object and perturbs object configurations in 3D while preserving demonstrated contact geometry. For mug racking, the mug and rack can move to new tabletop positions, but the approach and relative geometry must remain feasible. For trash sweeping, the trash and dustpan can be repositioned, but the sweep-to-container relation must survive.

The project page reports that 100 human demonstrations can be expanded with 500 additional DES episodes. The important part is not cosmetic data multiplication. The useful part is exposing the policy to object configurations it did not see in the original human demos, while keeping the manipulation geometry valid.

Installing the Open-Source Repo

The current repo expects Python 3.10, an editable LeRobot-style install, RLBench/LIBERO benchmark code, and SmolVLM2/SmolVLA model resources. Plan for a Linux machine with an NVIDIA GPU, matching CUDA/PyTorch versions, enough storage for datasets, and CoppeliaSim if you run RLBench.

The basic setup from the README is:

bash
sudo apt-get update
sudo apt-get install -y git git-lfs tmux xvfb xauth libegl1 libgl1-mesa-glx
git lfs install

cd /path/to/lerobot
conda create -n wepvla python=3.10 -y
conda activate wepvla
pip install -e ".[smolvla,libero]"

For the bundled RLBench implementation:

bash
pip install -e benchmarks/RLBench
python -c "import pyrep, rlbench; print('RLBench import ok')"

Before starting a training run, verify that PyTorch sees the GPU:

bash
python -c "import torch; print(torch.__version__, torch.cuda.is_available(), torch.cuda.device_count())"
nvidia-smi

Download the model resources:

bash
bash benchmarks/RLBench/scripts/download_vlm_models.sh

The expected layout is:

text
benchmarks/vlm_model/
├── SmolVLM2-500M-Video-Instruct/
└── smolvla_base/

For offline clusters, pass paths directly:

bash
VLM_MODEL_NAME=/path/to/SmolVLM2-500M-Video-Instruct \
VLM_WEIGHTS_PATH=/path/to/smolvla_base \
PYTHON=python \
bash benchmarks/RLBench/scripts/collect_data.sh \
  --dataset-root /path/to/rlbench_dataset

One small but expensive warning from the README: CUDA extensions must be rebuilt for the target machine's PyTorch and CUDA versions. Copying local build artifacts between machines is a fast way to create confusing import errors.

Training on RLBench

RLBench requires a compatible CoppeliaSim installation. You must pass COPPELIASIM_ROOT, LD_LIBRARY_PATH, and Qt-related variables to the commands. On a headless server, start Xvfb:

bash
Xvfb :99 -screen 0 1280x1024x24 -nolisten tcp >/tmp/rlbench-xvfb.log 2>&1 &

Collect data and build the cache:

bash
PYTHON=python \
DATASET_ROOT=/path/to/rlbench_dataset \
COPPELIASIM_ROOT=/path/to/CoppeliaSim \
LD_LIBRARY_PATH="/path/to/CoppeliaSim:${LD_LIBRARY_PATH:-}" \
QT_QPA_PLATFORM=xcb \
QT_QPA_PLATFORM_PLUGIN_PATH=/path/to/CoppeliaSim \
QT_PLUGIN_PATH="" \
bash benchmarks/RLBench/scripts/collect_data.sh \
  --dataset-root /path/to/rlbench_dataset

Train:

bash
DATASET_ROOT=/path/to/rlbench_dataset \
OUTPUT_ROOT=/path/to/rlbench_output \
GPU_IDS=0 \
bash benchmarks/RLBench/scripts/train.sh

Evaluate the 10 tasks:

bash
EVAL_POLICY_PATH=/path/to/checkpoint/pretrained_model \
EVAL_ROOT=/path/to/rlbench_eval \
EVAL_SAVE_VIDEO=0 \
EVAL_SAVE_ACTION_RECORDS=0 \
EVAL_SAVE_ACTION_CHUNKS=0 \
DISPLAY=:99 \
bash benchmarks/RLBench/scripts/evaluate.sh \
  --tasks close_box close_fridge close_laptop_lid phone_on_base stack_wine \
  sweep_to_dustpan take_frame_off_hanger \
  take_umbrella_out_of_umbrella_stand toilet_seat_down water_plants \
  --episodes 100

If you are new to RLBench, run one task and one episode first. Enable video and action logging while debugging. Only run the full evaluation after the simulator, display, dataset, and checkpoint path are all boringly correct.

Ten RLBench manipulation tasks used to evaluate WEPVLA — source: UMR project page
Ten RLBench manipulation tasks used to evaluate WEPVLA — source: UMR project page

Training and Inference on LIBERO

The LIBERO entry points live under benchmarks/song_real_libero/. The flow is similar: convert demonstrations, build the PointSeg cache, train, then evaluate.

Convert demonstrations:

bash
PYTHON_BIN=/path/to/python \
DEMO_ROOT=/path/to/libero_demos \
DATASET_ROOT=/path/to/libero_dataset \
bash benchmarks/song_real_libero/prepare_dataset.sh

Build the cache:

bash
PYTHON_BIN=/path/to/python \
DATASET_ROOT=/path/to/libero_dataset \
CACHE_ROOT=/path/to/libero_cache \
GPU_IDS=0 NPROC=1 \
bash benchmarks/song_real_libero/build_cache.sh

Train:

bash
PYTHON_BIN=/path/to/python \
DATASET_ROOT=/path/to/libero_dataset \
CACHE_ROOT=/path/to/libero_cache \
BASE_POLICY=/path/to/base_policy/pretrained_model \
OUTPUT_ROOT=/path/to/libero_output \
GPU_IDS=0 \
bash benchmarks/song_real_libero/train.sh

Evaluate:

bash
PYTHON_BIN=/path/to/python \
POLICY_PATH=/path/to/checkpoint/pretrained_model \
OUTPUT_DIR="benchmarks/song_real_libero/outputs/eval_$(date +%Y%m%d_%H%M%S)" \
CUDA_DEVICE=0 EPISODES=50 \
bash benchmarks/song_real_libero/evaluate.sh

At inference time, remember that this is not a plain image-to-action policy. The system builds a point-cloud observation, runs motion-aware segmentation, passes tokens through World and Ego streams, jointly denoises an action chunk, and executes the Ego action. For real deployment, log camera/depth health, point-cloud shapes, gripper state, predicted chunks, executed Ego trajectory, and task outcomes.

Four LIBERO suites used in the WEPVLA benchmark — source: UMR project page
Four LIBERO suites used in the WEPVLA benchmark — source: UMR project page

How to Read the Results

The headline numbers are strong, but they should be read in scope. On simulation, WEPVLA reports 85.7% mean success on the 10 RLBench tasks, compared with 82.3% for PointACT and 74.0% for HybridVLA in the project page table, while using a smaller 0.5B model. On LIBERO, WEPVLA reports 97.5%, compared with 96.0% for PointACT and 96.8% for OpenVLA-OFT.

The real-world comparison is more interesting. WEPVLA and HumanEgo are both trained with human-only data, about 10 minutes per task. WEPVLA reports 95.0% cross-environment success, 91.7% cross-embodiment success, 90.0% cross-setup success, and 91.7% average success across Tasks I-VI. HumanEgo reports 70.0%, 60.8%, 58.3%, and 60.8% under the same categories.

That does not mean a phone video plus 10 minutes of recording gives you a production robot. The reported system still depends on RGB-D sensing, point clouds, calibrated robot setups, task framing, and a controlled evaluation protocol. The useful lesson is narrower and stronger: if the action representation is geometric enough, human demonstrations become much more useful for robot policy learning.

When Should You Try UMR/WEPVLA?

UMR/WEPVLA is worth trying if your team has at least some of the following:

  • RGB-D or multi-view data, or the discipline to set up calibration properly.
  • Tabletop manipulation tasks with clear object motion: stacking, folding, sweeping, placing, racking, or closing.
  • A desire to compare human demonstrations with robot teleoperation, instead of assuming robot-only data is the only path.
  • Existing LeRobot, SmolVLA, or OpenVLA experiments that need better 3D grounding.
  • Enough GPU budget to run real training and evaluation, not just a notebook demo.

Be slower if the task requires heavy tactile feedback, unstable deformable-object handling, force-sensitive closed-loop control, or mobile manipulation with poor depth sensing. WEPVLA improves the representation. It does not remove the need for calibration, controller safety, dataset QA, or failure analysis.

Beginner Checklist

  1. Read the paper sections on UMR, WEPVLA, and DES before changing the code.
  2. Run evaluation with an available checkpoint before training your own model.
  3. Verify dataset paths, model paths, CUDA, and simulator imports separately.
  4. For RLBench, start with one task and one episode, with video/action logs enabled.
  5. For LIBERO, write each evaluation to a fresh output directory.
  6. When training a custom task, log both World Flow predictions and executed Ego actions.
  7. If results are poor, inspect point clouds, frame transforms, and gripper state before changing model hyperparameters.

UMR/WEPVLA matters because it reframes the policy question. Should a manipulation model learn object motion or robot action? UMR's answer is: learn both, force them to agree through geometry, and execute the robot-local part. That is a compact idea, backed by code, benchmarks, and enough practical detail for small labs to start testing it seriously.

Related Posts

  • OpenVLA deep dive
  • SmolVLA training in LeRobot
  • UMI/VLA: from human demos to robot policy
NT

Nguyễn Anh Tuấn

Robotics & AI Engineer. Building VnRobo — sharing knowledge about robot learning, VLA models, and automation.

Explore VnRobo

Fleet MonitoringROS 2 IntegrationAMR Solutions
← Previous
SimpleMemVLA: Native Video Memory for Long-Horizon VLA
Next →
RACE-VLA: 4x Longer Action Chunks

Related Posts

Tutorial
DSPv2: Dense Policy for Whole-Body Mobile Manipulation
vlawbcwhole-body
wholebody-vla

DSPv2: Dense Policy for Whole-Body Mobile Manipulation

DSPv2 fuses 3D point clouds with multi-view DINOv2 semantics for generalizable whole-body mobile manipulation — complete guide to setup, training, and inference.

9/15/202614 min read
NT
Research
SimpleMemVLA: Native Video Memory for Long-Horizon VLA
vlawhole-bodymanipulation
wholebody-vla

SimpleMemVLA: Native Video Memory for Long-Horizon VLA

SimpleMemVLA drops dedicated memory modules entirely, using Qwen3.5-4B's native video context as working memory — setting new SOTA on four memory-centric manipulation benchmarks.

9/20/202611 min read
NT
Tutorial
FluxVLA Engine: Hands-On Guide to the One-Stop VLA Platform
vlalerobotpi0
wholebody-vla

FluxVLA Engine: Hands-On Guide to the One-Stop VLA Platform

LimX Dynamics' open-source FluxVLA Engine covers the full pipeline from data to real-robot deployment, supporting Pi0, GR00T, and OpenVLA in one unified framework.

9/17/20269 min read
NT
VnRobo logoVnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam