RoboDojo tests GPT-6 Astra, finds a physics gap
Chris Paxton calls it an incredibly good robotics model, Chen Tessler says it has no business doing short-horizon execution.
The RoboDojo team ran a structured evaluation of GPT-6 Astra as an embodied agent, and posted the full report. Tianxing Chen shared it, listing what they tested: RoboDojo sim and real compared against GPT-5.5 and DeepSeek-Flash, humanoid high-level control, dexterous piano playing with RoboPianist, and a systematic study of in-context learning (the model picking up a new task from examples in its prompt rather than retraining).
The verdict is split. Astra shows strong semantic and spatial understanding and impressive in-context adaptation. But the team flags physical commonsense as the clear bottleneck, which they describe as a gap between understanding the world and reasoning about its physics. Chen credits @_wenbozhang as project lead on the full report and demos.
At the same time, physical commonsense remains a clear bottleneck
A third-attempt success on the snack tray
Andre Infante, who has been running his own outsider tests, says he tweaked the harness and hardware setup and got Astra to roughly succeed at the snack tray test on attempt 3. He is specific about the two failures before it. The first attempt was cancelled because of a system prompt issue. The second was killed because the robot dropped the OJ and could not reach it. His own read is that it works well, considering.
So nobody is claiming one-shot autonomy here. The person who ran the test is the one telling you it took setup changes and three tries.
The argument about where a general model belongs
Chris Paxton posted flatly that Astra is an incredibly good robotics model, and separately that the GPT of robotics will probably just be GPT. Chen Tessler pushed back. His position is that the model proves spatial understanding and reasoning can be solved by scaling up, and also that it makes zero sense for it to handle short-horizon execution, meaning the fast low-level control loop that actually moves the joints.
The skepticism is not only from researchers with a thesis. SkalskiP asked Paxton how his team actually uses it, saying his own zero-shot attempts at things he had seen online were mediocre at best. Paxton's answer was that most robotics benchmarks are pretty weak, which is less a defence of Astra than a shrug at the yardsticks everyone is using to judge it.
That is the shape of the week. One named evaluation with a named weakness, one honest real-hardware result that needed three runs, and a live disagreement about whether a general language model belongs anywhere near the control loop.
GPT-6 Astra is an incredibly good robotics model
GPT is both the moment for robotics (in the sense that it shows us we can pretty much solve spatial understanding and reasoning by scaling up), but at the same time it makes zero sense for it to handle short-horizon execution.


