Robotics Weekly

LIT trains robots to act first, then use vision

A two-stage recipe that keeps the action expert from latching onto backgrounds and camera angles.

Jiafei Duan has introduced Latent Interface Training, or LIT, a two-stage recipe for training VLAs (models that turn camera images and text into robot actions) and what the paper calls WAMs. His pitch is blunt. Your action expert may be learning vision to action shortcuts that hurt generalization once you step outside the training distribution.

Jiafei Duan
@DJiafei
X
Learn to act first, then learn how to use vision.
Sep 11, 2026 · View on X

How it works

Stage one trains the action expert with no images at all. It is conditioned on language, robot state, and the end-effector pose each action chunk should reach. Stage two adds vision, but only through a latent interface that is trained to reconstruct that same pose. The idea is that the model cannot lean on backgrounds, lighting, or a camera angle that happens to correlate with the right action, because vision only enters through a channel that carries pose information.

The numbers

On LIBERO-Plus, the authors report LIT lifting pi 0.5 from 68.97% to 79.67%, MolmoAct2 from 63.62% to 71.92%, FAST-WAM from 51.44% to 60.63%, and ImageWAM from 83.02% to 86.89%. So it is a recipe applied on top of existing models rather than a new model, and it helps the weaker ones most.

The real-world test is smaller but more concrete. MolmoAct2 trained on 300 demos across three tasks went from 53.3% to 70.0% under new lighting, 30.0% to 46.7% when given only the top camera, 50.0% to 63.3% with distractors added, and 74.7% to 88.0% in distribution. All of those are reported by the authors, not by an outside lab.

Shown versus claimed

This is a paper result, not a product demo. The gains sit in the 10 to 17 point range, which is real but nowhere near closing the gap that generalist foundation models have been claiming against tuned robot policies. Worth watching whether other groups reproduce it on their own stacks.

Get the next one by email

Robotics every day from the people building it. The demos, the deployments and the arguments worth your time. Every claim links back to the engineer.