Series
Ai Manipulation Agents
The "Ai Manipulation Agents" series has 5 parts — read them in order from part 1.
Manipulation
Multi-Agent Beats Single VLA Models | Manipulation Agents #1
ManiAgent hits 86.8% on SimplerEnv vs pi0 55.7% and CogACT 51.3% — zero robot fine-tuning. A deep-dive into the 3-agent architecture and why decomposition wins.

Perception Agent: Florence-v2 + AnyGrasp | Agents #2
ManiAgent's Perception Agent: Florence-v2 for zero-shot open-vocabulary detection, AnyGrasp for 6-DoF grasp pose generation. Build a complete Python module.

ALRM: Code-as-Policy vs Tool-as-Policy | Agents #3
ALRM's two execution modes: CaP generates Python to call robot APIs in one pass, TaP uses ReAct to loop per tool call. Benchmarked across 56 tasks, 10 LLMs.

Agentic Robot: SAP Protocol & Temporal Verifier
Run ds.py (DeepSeek-V3 subgoal decomposition) and main.py (OpenVLA on LIBERO). Implement a Temporal Verifier window — SAP protocol hits 79.6% LIBERO avg.

Sim-to-Real: Moving the SAP Pipeline from LIBERO to Robots
From 79.6% on LIBERO to real robots: calibrate camera-to-robot transforms, close domain gap via real-world augmentation, deploy SAP on Franka/UR5, under 100ms.