GPT-6 Astra hits 95% on a robot arm task, Chooi says
The model wrote end-effector poses directly, with no VLA in the loop, and researchers are already arguing about whether that is the cheap way to do it.
Jay Chooi posted that GPT-6 Astra scored 95% on a robot control task, up from 40% for Fable 5.1, with 6.2x fewer output tokens at 2.3x lower cost.
The Fable 5.1 number came from Chooi's own test days earlier. That one was a block into a bowl, 40% success against Fable 5's 5%, over 20 runs each, with 1.5x fewer tokens.
The part that got researchers talking is how the model did it. Jiafei Duan asked whether GPT-6 could just API call a MolmoAct2 checkpoint and ask it to put the block in the bowl. Chooi said no. GPT-6 used no VLA (a model that turns camera images and text into robot actions) at all, and specified the end-effector pose, meaning where the gripper should go, directly.
Real progress or an expensive detour
Chris Paxton called it "genuinely real progress here" while adding that there is a long way to go before these models can do control tasks.
A long way to go before these models can do control tasks, but genuinely real progress here
Max Fu pushed back on the framing. He argued that harnesses plus tool calls, things like IK (inverse kinematics, solving joint angles for a target gripper pose) and SAM3, get the same result far cheaper, faster and more reliably. His larger point is that self evolving harnesses and skill libraries will expand what these models can do, with VLAs as one component inside those libraries rather than the whole system.
Ken Goldberg picked up Fu's point and ran with it, saying it was nice to see Astra benchmarked on agentic robotics so quickly and that new research in harness design can improve performance further. He pointed to Graph as Policy, a paper he says was just accepted to CoRL.
The other GPT-6 robotics demo
Separately, lingxiao guo says GPT-6 did real2sim end to end, meaning rebuilding a real robot scene as a physics simulation. Guo says the inputs were only multi view RGB and robot actions plus one simple prompt, and the model calibrated the cameras, built the object assets, performed physics system identification, ran MuJoCo and rendered in Blender.
Shown versus claimed. The 95% post does not spell out the task setup, and the earlier Fable comparison ran only 20 trials per model, so these are individual researcher tests, not a standardised benchmark. What is not in dispute is the reflex. A frontier chat model shipped and robotics people had it driving an arm within days, then argued about harness design instead of whether the idea made sense at all.
harnesses + tool calls (IK, SAM3, etc.) can achieve the same result far cheaper, faster, and more reliable.
Here GPT-6 didn't use any VLAs at all and specify the EEF pose directly.


