Astra 35%, Fable 15%, human teleop 100% on Pandroid robots
A controlled real-hardware benchmark put two general models on robot control, and Chris Paxton says robotics labs should be asking why it was not their models.
Pantograph put GPT-6 Astra and Claude Fable 5.1 on its Pandroid robots and ran a controlled benchmark across eight manipulation tasks. Astra completed 35% of attempts, Fable 15%, with p roughly 0.001. Human teleoperators, meaning people driving the robot remotely, completed every task.
Astra completed 35% of attempts, Fable 15% (p≈0.001). Human teleoperators completed every task.
That last number is the one people latched onto. @addadryanw said the teleop gap is where the attention belongs, and that 35% and 15% both say contact-rich manipulation is still unsolved. Chris Paxton agreed, writing that none of the results imply contact-rich manipulation is anywhere close to solved.
Why Paxton called it soul searching
Paxton's first reaction was not about the scores. It was about whose models were on the test bench. He called it a real-world test of robot control and noted, tellingly, that it was being done with Fable and Astra rather than any robotics foundation model, then said it was time for more soul searching on the part of his field.
That is the sting. Purpose-built robot foundation models are the thing a large slice of the field has spent the last few years building, and the group running a head-to-head hardware comparison reached for two general-purpose models instead.
What the numbers do and do not say
Shown versus claimed matters here. Pantograph reported completion rates on its own robots and its own eight tasks, with a stated significance level, so the Astra-over-Fable gap is a measured result rather than a vibe. Nobody outside Pantograph has replicated it. The task list itself has not been published in the posts, so how hard these eight tasks are relative to anything else is not something the sources settle.
Astra has been in a run of outside tests lately, including a thread claiming it beat MolmoAct2 46% to 12% on two-armed tasks. This one is different in that it puts a human baseline in the same column, and the human baseline is a clean 100%.
Paxton added one more read in a follow-up. Asked about SLAM, he said he is not sure, and that in a lot of ways what these models are solving is local spatial reasoning, which is enough for robotics in the sense that they can do pick and place tasks. That squares with the numbers. A third of attempts finished is real progress for a model that was not built to drive a robot, and it is nowhere near a human with a controller.
time for some more soul searching on the part of my field i think
I’d focus on the teleop gap. 35% and 15% both say contact-rich manipulation is still unsolved.


