VnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
VnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam
VnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
  1. Home
  2. Blog
  3. Run GALA on RoboCasa-GR1
wholebody-vlagalavlalatent-actionrobocasa-gr1humanoidmanipulation

Run GALA on RoboCasa-GR1

A practical GALA guide for geometry-aware latent actions, multi-embodiment VLA training, and RoboCasa-GR1 inference.

Nguyễn Anh TuấnSeptember 23, 202614 min read
Run GALA on RoboCasa-GR1

GALA, short for Geometry-Aware Latent Action Modeling, is one of the more useful papers to study if you care about VLA policies that must learn from multiple robot bodies. The core question is not simply "can we train a larger VLM?" It is more basic: when a human hand, a dexterous robot hand, and a parallel-jaw gripper all use different action spaces, can we learn a shared latent action that still preserves wrist motion, finger articulation, and end-effector geometry?

The original paper is GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments. The project page is puzhenyuan.github.io/GALA-website, the official repository is PuzhenYuan/GALA, and the RoboCasa-GR1 checkpoint is hosted on Hugging Face. This guide is written for builders: we will cover the paper idea, the architecture, installation, RoboCasa-GR1 inference, the training blueprint, and the results that matter.

If you are new to VLA systems, read What are VLA models? first. GALA sits one layer deeper. It does not only ask what action the policy should output; it asks what action representation can transfer across embodiments without erasing fine motor detail.

Tool recommendations

VLA train/deploy stack

Train on cloud/workstation, then deploy optimized models to Jetson or the robot computer.

Cloud GPU for VLA / policy training Use for imitation learning, diffusion policies, RL, and robotics model fine-tuning. View cloud GPU → NVIDIA Jetson Orin NX / Orin Nano Edge deployment hardware for perception, logging, and optimized inference. View Jetson → Hugging Face / robotics dataset hosting Host datasets, checkpoints, and model cards for cleaner LeRobot/VLA workflows. View platform →

GALA teaser: geometry-aware latent actions combine RGB transitions and 3D end-effector motion - source: PuzhenYuan/GALA repo
GALA teaser: geometry-aware latent actions combine RGB transitions and 3D end-effector motion - source: PuzhenYuan/GALA repo

The Problem GALA Solves

In conventional imitation learning, the action label is usually the robot's native control vector: joint positions, delta end-effector pose, gripper command, finger joints, waist motion, or another dataset-specific command format. That works when all data comes from one robot. It becomes brittle when VLA pretraining tries to combine many sources: human manipulation videos, simple gripper robots, dexterous robot hands, and humanoids such as GR-1.

The mismatch is severe:

  • Human videos often have no robot action labels.
  • Parallel-jaw grippers have far fewer DoF than dexterous hands.
  • Dexterous hands require finger-level articulation, which is hard to align directly with human hands.
  • Humanoid tabletop manipulation may include two hands, waist motion, cameras, and timing that differ from fixed-arm datasets.

Latent Action Models, or LAMs, are one way around this. Instead of supervising a policy with native robot actions, a LAM learns discrete latent tokens from visual start-goal transitions. Those tokens describe what happened between two observations, so they can be learned even from action-free videos. The weakness, as GALA points out, is that image-only LAMs often capture scene-level change but miss fine-grained end-effector articulation. An image transition may show that an object moved, while hiding how the fingers curled, how the wrist rotated, or how the hand geometry changed during contact.

GALA adds a geometric branch. For each start-goal interval, the system builds 3D point clouds of the left and right end-effectors. For human hands, the paper uses WiLoR to estimate MANO meshes and keypoints from RGB frames. For robot hands and grippers, the point clouds come from URDF or MJCF geometry transformed by forward kinematics using the recorded joint configuration. The result is a latent action that does not only encode object displacement; it also encodes how the end-effector moved in 3D.

The Core Idea

Imagine two frames: before interaction and after interaction. A standard image-based LAM looks at the RGB transition and compresses the difference into a discrete token. GALA keeps that visual token but adds point-cloud transitions of the end-effectors. All point clouds are expressed in a wrist- or root-centered local frame with a common axis convention. The representation does not require point correspondence, joint correspondence, or shared mesh topology across embodiments.

That design choice is important. If human MANO hands, Fourier hands, XHand, and Robotiq grippers all had to share the same mesh or joint indexing, the system would be fragile. GALA avoids that trap by treating each end-effector as an unordered surface point cloud.

But simply adding point clouds is not enough. A naive point-cloud representation can become too embodiment-specific: it may memorize morphology instead of learning motion semantics. GALA addresses this with Unified End-effector Motion Representation, or UEMR. UEMR has three main pieces:

  1. Unified bimanual motion latent: the geometric latent jointly represents the bimanual transition instead of assigning rigidly separate latents to left and right hands. For single-arm gripper data, the same point cloud can feed both branches to keep the interface consistent.
  2. Pair-consistent geometric augmentation: the same random 3D transform is applied to both the start and goal point clouds. Absolute coordinates change, but relative motion is preserved, which discourages coordinate shortcuts.
  3. Bidirectional transition learning: the model learns both forward and backward temporal transitions with shared encoders, codebooks, and decoders. This gives the latent more supervision for relative dynamics.

In plain terms, GALA trains the latent action to answer: "what did the end-effector do in 3D?", not "which robot hand is this?".

Architecture: Two Training Stages

GALA Stage 1: visual encoder, point-cloud encoder, UEMR, and latent action reconstruction - source: GALA project page
GALA Stage 1: visual encoder, point-cloud encoder, UEMR, and latent action reconstruction - source: GALA project page

Stage 1: Geometry-Aware Latent Action Learning. The model receives start-goal RGB observations, a language instruction, and start-goal point-cloud pairs for the left and right end-effectors. RGB observations are encoded with frozen DINOv2 and SigLIP. Hand point clouds are encoded with a frozen Point Transformer V3. The visual and geometric features are fused with language through a shared spatiotemporal Transformer.

The output is split into two complementary token streams:

  • Visual latent action tokens, which model scene-level dynamics.
  • Geometric latent action tokens, which model bimanual 3D end-effector transitions.

The geometric tokens are quantized with a shared EMA-VQ codebook. A geometry decoder receives the initial geometry and the quantized latent, then reconstructs the goal point cloud. The reconstruction is supervised with Chamfer Distance. Because the decoder is conditioned on the initial geometry, the latent does not need to store static morphology; it is encouraged to store the start-goal transition.

GALA Stage 2: frozen latent action model supervises VLM bridge tokens and a shared DiT action expert - source: GALA project page
GALA Stage 2: frozen latent action model supervises VLM bridge tokens and a shared DiT action expert - source: GALA project page

Stage 2: Geometry-Aware VLA Co-Training. Once Stage 1 is trained, the latent-action model is frozen. The VLA backbone receives the current observation and instruction, plus two groups of learnable bridge tokens. One group predicts visual latent-action codes. The other predicts geometric latent-action codes. Both are trained with cross-entropy against the discrete codes produced by the frozen GALA model.

The real action side is deliberately not forced into one artificial unified action vector. For human demonstrations, the paper defines actions as wrist 3D positions and rotations in the camera frame plus finger keypoint positions in the local wrist frame. For robot embodiments, it keeps each dataset's native action space. To handle these differences, the action expert uses a shared diffusion transformer (DiT) plus lightweight embodiment-specific action heads. The shared DiT learns transferable visuomotor dynamics, while each head maps the shared representation to the correct action dimension and control semantics. The final objective combines latent prediction losses with a flow-matching objective for continuous action generation.

Hardware and Environment

The official repository expects Linux, Python 3.10, an NVIDIA GPU with CUDA 12.4 support, a CUDA build toolchain, and EGL for headless MuJoCo rendering. If you are a beginner, do not start with the full 8-GPU evaluation. First run one task, one episode, and one GPU. Once the pipeline works, scale up.

A reasonable setup looks like this:

OS: Ubuntu 22.04 or similar
GPU: NVIDIA GPU with CUDA 12.4 support
Python: 3.10
Renderer: MUJOCO_GL=egl
Disk: enough for the GALA checkpoint, Qwen2.5-VL-3B, and RoboCasa assets

The README setup flow is:

git clone [email protected]:PuzhenYuan/GALA.git
cd GALA
conda create -n gala python=3.10 -y
conda activate gala
bash examples/environment_setup.sh

The setup script installs PyTorch 2.5.1 with CUDA 12.4, project requirements, flash-attn==2.7.1.post4, fixed revisions of robosuite and robocasa-gr1-tabletop-tasks, a RoboCasa patch, editable packages, tabletop assets, and then runs pip check. If flash-attn fails, check the CUDA compiler/toolchain. If MuJoCo rendering fails, check MUJOCO_GL=egl, NVIDIA drivers, and EGL libraries in your host or container.

Download Checkpoints

The public RoboCasa checkpoint is ypz21/GALA_robocasa_gr1. The README notes that the model repository may require an account with access, so authenticate with Hugging Face first:

hf auth login
hf download ypz21/GALA_robocasa_gr1 --local-dir checkpoints/checkpoint_robocasa_gr1
hf download Qwen/Qwen2.5-VL-3B-Instruct --local-dir checkpoints/Qwen2.5-VL-3B-Instruct

The RoboCasa-GR1 checkpoint uses Qwen2.5-VL-3B-Instruct as the backbone. During evaluation, set GALA_BACKBONE_PATH to the local backbone directory so your run does not depend on network downloads.

Smoke-Test Inference

Before launching 24 tasks with 50 episodes each, run a tiny smoke test:

mkdir -p outputs

PYTHON_BIN="$(command -v python)" \
GALA_BACKBONE_PATH="$PWD/checkpoints/Qwen2.5-VL-3B-Instruct" \
GPU_IDS=0 \
PROCS_PER_GPU=1 \
PORT_BASE=5810 \
N_ENVS=1 \
N_EPISODES=1 \
EVAL_MAX_TASKS=1 \
EVAL_TAG=_smoke \
DATA_CONFIG=fourier_gr1_arms_waist_gausNorm_crop_cam_ego_joints_only \
bash examples/run_eval_parallel.sh checkpoints/checkpoint_robocasa_gr1 id

This checks the entire path: load checkpoint, start the inference service, create a RoboCasa environment, render headlessly, query the policy, write results, and save video. If the smoke test fails, inspect the server/client logs under outputs/evaluation_sim_id_1envs_smoke/. Common issues are missing RoboCasa assets, wrong checkpoint path, occupied ports, and broken EGL rendering.

When the smoke test is stable, run the full evaluation from the README:

mkdir -p outputs
set -o pipefail

PYTHON_BIN="$(command -v python)" \
GALA_BACKBONE_PATH="$PWD/checkpoints/Qwen2.5-VL-3B-Instruct" \
GPU_IDS=0,1,2,3,4,5,6,7 \
PROCS_PER_GPU=3 \
PORT_BASE=5810 \
N_ENVS=1 \
N_EPISODES=50 \
EVAL_TAG=_gala_parallel \
DATA_CONFIG=fourier_gr1_arms_waist_gausNorm_crop_cam_ego_joints_only \
bash examples/run_eval_parallel.sh checkpoints/checkpoint_robocasa_gr1 id \
2>&1 | tee outputs/eval_gala_robocasa_gr1_id.log

This evaluates 24 ID tasks, 50 episodes per task, using 8 GPUs and 3 workers per GPU. Outputs are saved under outputs/evaluation_sim_id_1envs_gala_parallel/, including results.json, client/server logs, and rollout videos.

RoboCasa-GR1 demo: Cup to drawer + close, source: GALA project page

Training Blueprint

The public repository currently focuses on setup and checkpoint evaluation. So this section should not be read as a copy-paste full training recipe. It is a faithful blueprint of the paper pipeline, useful if you want to reproduce GALA once full training code and data preprocessing are available, or if you want to implement a similar pipeline in your own stack.

The training process has three layers.

Layer 1: Build geometry inputs. For human videos, run a hand estimator such as WiLoR to obtain MANO meshes and keypoints, align each hand to a palm/wrist local frame, and sample point clouds from the hand surface. For robot data, use the URDF or MJCF model plus recorded joint states to forward-kinematically transform end-effector link surface samples, aggregate them, and resample to a fixed number of points. Missing hands are handled with validity masks.

Layer 2: Train the geometry-aware LAM. Sample a start-goal interval. Encode RGB frames with frozen DINOv2 and SigLIP. Encode point clouds with frozen PTv3. Fuse the visual, geometric, and language features through a spatiotemporal Transformer. Quantize geometric tokens with EMA-VQ. Decode the goal geometry from the initial geometry plus the latent code. Optimize the visual VQ objective, Chamfer Distance for UEMR geometry reconstruction, EMA-VQ regularization, and both forward/backward transition objectives.

Layer 3: Co-train the VLA policy. Freeze the GALA latent-action model. For each training sample, generate visual and geometric latent code targets. Train bridge tokens in the VLM to predict those codes with cross-entropy. In parallel, train the shared DiT action expert with a flow-matching action loss. For each sample, only the matching embodiment-specific action head is activated and optimized.

The practical checklist is:

1. Normalize datasets into episodes with RGB, instructions, state, and actions when available.
2. Generate end-effector point clouds for every usable timestep.
3. Sample start-goal intervals and build validity masks.
4. Train Stage 1: visual latent branch plus UEMR geometry branch.
5. Freeze the latent-action model and encode latent code targets.
6. Train Stage 2: bridge-token CE losses plus flow-matching action loss.
7. Evaluate on RoboCasa-GR1 and, if possible, real-world dexterous tasks.

If you already know LeRobot or OpenVLA, the biggest conceptual difference is that GALA does not convert all actions into a single shared vector. It keeps native action spaces for deployment, but uses shared latent-action supervision for transfer. That pattern matters for whole-body and humanoid VLA because forcing every body into one artificial action schema often discards useful control semantics.

Results That Matter

RoboCasa-GR1 rollout image for Cup to drawer + close - source: GALA project page
RoboCasa-GR1 rollout image for Cup to drawer + close - source: GALA project page

The paper evaluates GALA at three levels: latent motion probing, cross-embodiment retrieval, and downstream VLA performance.

For fine-grained motion probing, the latent encoders are frozen and a lightweight MLP predicts end-effector translation, wrist rotation, and finger/gripper articulation on a held-out XHand split. GALA gets the best reported errors: 5.359 cm position error, 7.898 degrees rotation error, and 7.585 degrees finger error. This supports the claim that the geometric latent retains fine motor information better than image-only latent actions.

For cross-embodiment retrieval, the benchmark annotates grasp, hold, and release transitions across a human hand, a parallel-jaw gripper, and two dexterous robot hands. Retrieval is strictly cross embodiment. GALA reaches R@1 of 45.67 on the human-robot track and 46.91 on the robot-only track, outperforming METIS, native kinematics, OPFA, and the no-UEMR ablation. This test is valuable because it asks whether the latent groups motion type across different bodies, not just within the same robot.

For RoboCasa-GR1, the paper reports two settings. In GR-1-only training, GALA reaches 55.7% success, above UniVLA at 48.0, METIS at 43.8, native kinematics at 51.8, and OPFA at 53.5. In multi-embodiment co-training with Fourier-hand, XHand, human-hand, and Robotiq gripper data, GALA reaches 68.3% average success across 24 tasks. That is higher than UniT at 66.8 and JoyAI-RA at 63.2 in the paper's comparison table. Removing UEMR drops the result to 58.6, which makes UEMR a central contribution rather than a cosmetic addition.

Setting Method Success
GR-1 only UniVLA 48.0%
GR-1 only OPFA 53.5%
GR-1 only GALA 55.7%
Multi-embodiment UniVLA 53.6%
Multi-embodiment GALA w/o UEMR 58.6%
Multi-embodiment GALA 68.3%

On real-world XHand tasks, GALA reaches 75.5% average success across Pick, Push, Press, and Flip. It beats HARP-VLA at 71.5 and GALA without UEMR at 68.5. The Press Button and Flip Cup tasks are especially relevant because they require precise contact and finger-level control, which is exactly what the geometric latent branch is designed to preserve.

When Should You Use GALA?

GALA is a good fit when you need one of these:

  • You want to evaluate the public RoboCasa-GR1 checkpoint.
  • You are researching cross-embodiment VLA and need action supervision that is less tied to native action spaces.
  • You want to use human videos or heterogeneous robot datasets while preserving fine-grained hand motion.

It may be too heavy if you only need a quick single-robot gripper baseline for simple pick-and-place tasks. In that case, behavior cloning, diffusion policy, or a lighter 3D policy can be faster to debug. But once you move toward humanoid manipulation, dual-hand dexterity, or multi-embodiment pretraining, UEMR is a pattern worth studying.

Beginner Debug Checklist

Debug in this order:

1. Confirm Python is 3.10.
2. Confirm PyTorch sees CUDA.
3. Confirm flash-attn imports.
4. Confirm RoboCasa tabletop assets are downloaded.
5. Confirm MUJOCO_GL=egl works.
6. Confirm the GALA checkpoint and Qwen backbone paths are correct.
7. Run a smoke test with N_EPISODES=1 and EVAL_MAX_TASKS=1.
8. Inspect results.json and rollout videos before full evaluation.

Do not ignore the videos. A success rate only tells you whether the final condition passed. The rollout video tells you whether the real failure is perception, timing, collision, gripper closure, action chunking, or rendering.

Takeaway

GALA is a strong example of the next wave of VLA work: improving not only the model size, but the action representation. By combining visual latent actions with geometric latent actions, then using UEMR to preserve motion semantics across human hands, robot hands, and grippers, GALA makes multi-embodiment pretraining less dependent on handcrafted action alignment. The 68.3% RoboCasa-GR1 result suggests the idea is not just elegant; it improves downstream policy performance.

If you are building a VLA stack for humanoids or dexterous robots, treat GALA as a design pattern: keep native actions for execution, but learn a shared latent action space for pretraining and transfer.

Related Posts

  • OpenVLA deep dive
  • LeRobot ecosystem guide
  • Fine-tune GalaxeaVLA G0.5
NT

Nguyễn Anh Tuấn

Robotics & AI Engineer. Building VnRobo — sharing knowledge about robot learning, VLA models, and automation.

Khám phá VnRobo

Fleet MonitoringROS 2 IntegrationAMR Solutions

Related Posts

NEWResearch
Opt2VLA: Dạy VLA dự đoán lực tiếp xúc cho humanoid
vlawholebody-vlahumanoid
wholebody-vla

Opt2VLA: Dạy VLA dự đoán lực tiếp xúc cho humanoid

Hướng dẫn Opt2VLA — kết hợp trajectory optimization, RL controller và GR00T-N1.7 để humanoid Digit thực hiện whole-body manipulation với kiểm soát lực chính xác từ ngôn ngữ.

9/28/202611 min read
NT
Research
LeVERB: Điều khiển toàn thân humanoid bằng ngôn ngữ-thị giác tiềm ẩn
wholebody-vlahumanoidvla
wholebody-vla

LeVERB: Điều khiển toàn thân humanoid bằng ngôn ngữ-thị giác tiềm ẩn

LeVERB (UC Berkeley) — framework phân cấp đầu tiên cho điều khiển toàn thân humanoid bằng latent VLA, zero-shot sim-to-real trên Unitree G1, đạt 58.5% thành công.

6/24/202613 min read
NT
Research
ROVE: Human Intervention làm RL Signal cho VLA Humanoid
rovevlareinforcement-learning
wholebody-vla

ROVE: Human Intervention làm RL Signal cho VLA Humanoid

ROVE dùng Optimistic Value Estimation (OVE) để fine-tune VLA humanoid manipulation từ human intervention imperfect — pipeline thực tế từ XPENG Robotics.

6/22/202612 min read
NT
VnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam