Robotics Weekly

Lerrel Pinto sees "deep despair" as Astra zero-shots benchmarks

Chris Paxton started it, then walked part of it back within minutes.

Chris Paxton set it off with one line. "If you are a robot foundation model guy and not a deployment guy this must be really concerning", he posted. Lerrel Pinto followed a day later, saying he is sensing deep despair in academics over the past week, and naming Astra, Fable and Muse as models zero-shotting benchmarks in robotics and world models. He called it an uneasy pill to swallow, and said this is what step jumps in progress look like.

Chris Paxton
@chris_j_paxton
X
If you are a robot foundation model guy and not a deployment guy this must be really concerning
Sep 11, 2026 · View on X

The replies came from people who have been in the field a long time. Ken Goldberg told Pinto he hears him, but that it is also thrilling to witness a paradigm shift in robotics. Michael Black told him the field has been here before and to treat it as an opportunity to jump ahead, pointing out that there are many unsolved problems and now new powerful tools that may help solve them. Mahi Shafiullah asked whether this is a GPT moment in robotics, and said his team's MolmoSpaces benchmark seems to suggest so.

No numbers on the table

Nobody in these posts put a score down. No trial counts, no task list, no head to head against a named robotics VLA (a model that turns camera images and text into robot actions). Paxton walked his own line back within minutes, saying he actually thinks this is more because most robotics baselines are not so good and are very contact light. A lot of pick and place, he added, you can do pretty well with segment anything plus a very light language model. So the despair is real, but the evidence on show says as much about how thin robot baselines are as it does about the big models.

What academics should do instead

Black's longer post argues the problem is structural. He says any paper you see at a conference is likely two years out of date, and in AI today two years means your work is likely irrelevant. His fix is that every project should start by trying really hard to solve the problem with existing tools, and every paper should open with a detailed experimental analysis of how existing models perform and why they fail. He wants a section like related work, but for current large models. He also made this personal, describing his own ECCV paper VIGA as fully out of date before he presented it, which we covered earlier.

Jeannette Bohg agreed with much of it and went further on the conference side. She said she has been bored by the main track of CoRL for years and already knows 90% of the works, and that her favourite part is the workshops, where people present what they just submitted. Her question was whether conferences should be all workshops.

Lerrel Pinto
@LerrelPinto
X
I’m sensing deep despair in academics over the past week. Astra, Fable, Muse are zero-shotting benchmarks in robotics
Sep 12, 2026 · View on X
Chris Paxton
@chris_j_paxton
X
most of the robotics baseline are not so good
Sep 11, 2026 · View on X

Get the next one by email

Robotics every day from the people building it. The demos, the deployments and the arguments worth your time. Every claim links back to the engineer.