Robotics Weekly

Flex-pi world action model predicts 3D, not just video

A 6 billion parameter model adds geometry and object semantics to video prediction, and the authors say it needs far fewer demos.

Flex-pi is a world action model (a model that learns to predict what the world will look like next while also learning to act in it) that predicts 3D pointmaps and DINO features alongside color video, instead of only reconstructing images. Ge Yan (@GeYan_21), Jesse Zhang (@Jesse_Y_Zhang) and team trained it at 6 billion parameters. It gets a full walkthrough in episode 106 of RoboPapers, hosted by @DJiafei and @ruijie_sg.

Chris Paxton
@chris_j_paxton
X
it becomes extremely data efficient, and is able to generalize to really complex, precise long horizon tasks
Sep 23, 2026 · View on X

Why geometry instead of pixels

RoboPapers frames the problem simply. World action models are getting popular because they learn to predict the world jointly with learning how to act on it, but those predictions are usually built purely on reconstructing color images from video. The account argues that is limiting, because color is far from the most important quality for a robot moving through the world. Geometry and object semantics matter more.

RoboPapers
@RoboPapers
X
color is far from the most important quality for a robot moving around in the world
Sep 23, 2026 · View on X

So Flex-pi predicts those directly. Pointmaps give the model an explicit sense of where surfaces sit in 3D. DINO features carry object identity, the kind of signal that tells you this is the same mug from a different angle rather than a new patch of pixels. Both are predicted along with color, not instead of it.

What the authors claim

The payoff, according to RoboPapers, is a policy that is much more demonstration-efficient, generalizes well, and can perform complex long-horizon tasks. Chris Paxton, who hosts the podcast, puts it more strongly, saying the model becomes extremely data efficient and can generalize to really complex, precise long horizon tasks. He also added an auto-generated transcript to the Substack page for the episode, with the caveat that it may not be perfect.

Shown versus claimed

The data-efficiency and generalization claims come from the authors and the podcast writeup. Neither post lists comparison numbers against other world action models, and no outside evaluation has been posted. So the interesting part right now is the design choice rather than a verified score. If predicting geometry and semantics really does cut the number of demonstrations a policy needs, that is the kind of result other labs will try to reproduce fast, because demo collection is the expensive part of every manipulation pipeline.

What would settle it is a head to head against a color-only world action model on the same tasks with the same demo budget. Nobody has posted that yet.

Get the next one by email

Robotics every day from the people building it. The demos, the deployments and the arguments worth your time. Every claim links back to the engineer.