Series
Robot 3d Manipulation Vla
The "Robot 3d Manipulation Vla" series has 7 parts — read them in order from part 1.
Research & Trends
Why 2D RGB VLA Is Not Enough for Manipulation
Robot policies and VLAs trained on 2D RGB images hit a hard ceiling. This opening article explains the 5 core limitations of 2D perception and why 3D-aware manipulation is the inevitable next step.

Robo3R: Feed-Forward 3D Reconstruction for Manipulation
From ordinary RGB cameras and robot state, Robo3R builds metric-scale 3D geometry in real time within the robot frame — no depth sensor, no SLAM, 43 Hz on single view.

DP3: Point-Cloud 3D Diffusion Policy Hands-On
DP3 turns a point cloud into a direct input for a diffusion policy — 24.2% relative improvement across 72 tasks, 85% success on real robots. Hands-on: install, train, eval from the YanjieZe/3D-Diffusion-Policy repo.

3D Perception for Whole-Body Humanoids: Beyond-FOV
A humanoid that walks and reaches far needs perception beyond the camera FOV plus active spatial reasoning. Analyzing Omni-Manip and Active Spatial Reasoning — two 2026 humanoid papers.

WholeBodyVLA: Action-Free Video to Loco-Manipulation
WholeBodyVLA learns from action-free egocentric video and connects a high-level VLA to a low-level RL controller for whole-body loco-manipulation. Analyzing the ICLR 2026 paper, tested on AgiBot X2.

Data Collection: Teleop vs Robot-Free vs Egocentric Video
Three strategies for whole-body manipulation data — teleop, robot-free demos (HuMI), and action-free egocentric video. A cost/quality/throughput comparison and how to choose by budget.

A 3D Manipulation Roadmap for the Unitree G1 Humanoid
The capstone: combine Robo3R, DP3, Omni-Manip, and WholeBodyVLA into a concrete deployment pipeline for the Unitree G1 — from sensor placement, policy choice, and data strategy to a sim2real checklist.