Series
Vima Humanoid Manip
The "Vima Humanoid Manip" series has 5 parts — read them in order from part 1.
Manipulation
VIMA Architecture: Cross-Attention Transformer Explained
Decode VIMA's cross-attention architecture: how T5 encodes multimodal prompts and XAttn GPT generates motor commands for 17 robot manipulation tasks.

VimaBench: 17 Tasks and 4-Level Generalization Protocol
Install VimaBench, run the 200M checkpoint demo, and analyze all 4 eval partitions — from placement to novel task generalization. Which level matters most for humanoid deployment?

Object Tokenizer: From Raw Pixels to Object Tokens with Mask R-CNN + ViT
A deep dive into VIMA's object tokenizer: Mask R-CNN objects, ViT crop features, bbox MLP position signals, and T5 prompt assembly.

Dataset 650K: Collecting Large-Scale Multi-Task Manipulation Data for VIMA
Deep dive into VIMA's 650K trajectory dataset: procedural generation in PyBullet, pkl file format, HuggingFace download, and why just 1% of the data beats all baselines.

Humanoid Adaptation: Extending VIMA to High-DoF Humanoid Robots
VIMA series capstone: adapt from tabletop 6-DoF to Unitree G1 23-DoF — freeze backbone, swap action head, bridge data gap with teacher-student and LoRA fine-tuning.