All posts
Page 2 of 18

MIT Han Lab's VLASH makes VLAs real-time via future-state-aware asynchronous inference — 11.8× lower reaction latency, no architecture changes, open-source.

Complete guide to training SONIC whole-body controller with GR00T-WholeBodyControl open-source on 100M+ motion capture frames, deployed on Unitree G1.

VLAC integrates a learned critic into the real-world RL loop, lifting success rate from 30% to 90% in just 200 episodes — no reward engineering required.

CoorDex uses body and hand latent priors to train Unitree G1 for continuous walk-grasp-carry without stopping. Open-source code from UNC Chapel Hill.

Deep dive into DyPES-VLA — cross-embodiment SOTA: 98% LIBERO, 89% RoboTwin 2.0, 75.6% real-world across 3 robots via shared Dynamics Prior and MoE Action Head.

A practical guide to SLIM-0.5B, a compact VLA manipulation policy with action-grounded predictive latents on LIBERO.

A practical guide to πR² for GR00T-N1.7: architecture, setup, fine-tuning, 25 Hz inference, and manipulation results.

A guide to ω-0's 3-stage architecture — the latent predictive WAM hitting 81.8% success on 11 household tasks, letting humanoids walk and grasp at once.

Run MiniVLA on LIBERO/MuJoCo, add perturbation tests, and understand how to report a 97.92% robustness success rate.

A practical World-to-Wrist VLA guide: wrist-view generation, wrist CoT labels, LeRobot data setup, training, inference, and results.

Technical guide to Qwen-VLA from Alibaba: Qwen3.5-4B + DiT decoder architecture, 4-stage training pipeline, 97.9% on LIBERO, beating pi0.5 on real-world ALOHA.

FM-VLA compresses force-sensor history into 8 VAE tokens, helping robots count button presses, wipe dishes, find hidden objects — 83.3% success, 3.3ms overhead.