Robotics Weekly

Grounded Action Model beats pi-0.5 when the target changes

A robot policy built on a frozen 3D grounding backbone posts 61% average success on LIBERO-PRO, and the authors still call grounding errors an open problem.

Jiafei Duan's group has introduced GAM, the Grounded Action Model, a robot policy that starts from a pretrained 3D grounding model instead of a language or video model. The grounding backbone stays frozen. Only the action head is trained.

Jiafei Duan
@DJiafei
X
Ground first. Then learn to act.
Sep 22, 2026 · View on X
Jiafei Duan
@DJiafei
X
GAM is a starting point. Grounding errors and omitted context, such as unselected obstacles, remain challenges.
Sep 22, 2026 · View on X

The pitch is in one line from the thread. Ground first, then learn to act. In practice, a pretrained 3D detector, WildDet3D in their setup, localizes the task objects and estimates their geometry. Image tokens keep features around the selected objects and the robot arm, detection tokens carry the object point clouds and metric geometry, and an MM-DiT combines those with robot state history and language to produce action chunks.

The numbers

On LIBERO-PRO, a simulated benchmark, GAM reports 61% average success across 16 settings against 53% for pi-0.5. The gap that people will argue about comes from LIBERO-Spatial. When the target object changes, GAM reports 88% and pi-0.5 reports 1%. The objects and the actions are familiar in that test. What changes is which object the robot is told to act on.

On RoboTwin 2.0 across 50 tasks, GAM reports 55.3% average success against 52.0% for Spatial Forcing, and 47.6% on randomized scenes against 30.4% for Abot-M0. There, action policies are trained per task on clean demos while the grounding backbone is adapted to simulator images separately and then frozen.

Targets can be specified by language, by clicking a point, or by drawing a box, which the group argues cuts language ambiguity and gives a higher level vision language model a direct handle on the policy. For long horizon and memory dependent tasks, a Molmo2 planner picks targets using current observations and episode history, and GAM does the acting. On a Franka arm, GAM plus Molmo2 reaches 64.7% in distribution and 49.8% out of distribution step completion.

Shown versus claimed

These are the authors' own numbers, mostly on simulated benchmarks, not an outside evaluation. The thread itself calls GAM a starting point and lists grounding errors and omitted context, such as obstacles that were never selected, as remaining challenges. Paper and code links are in the thread. The work is led by @gehao_zhang_ with @weikaih04, @Shailes_h_ and @RanjayKrishna.

Get the next one by email

Robotics every day from the people building it. The demos, the deployments and the arguments worth your time. Every claim links back to the engineer.