GPT-6 Astra beats MolmoAct2 46% to 12%, thread says
The first head-to-head number against a named robotics VLA, and the field's first reaction was to argue about the baseline.
Jay Chooi (@chooi_jeq) posted a thread putting GPT-6 Astra at 46% against 12% for MolmoAct2, which the thread calls a state-of-the-art robotics VLA (a model that turns camera images and text into robot actions). The comparison covers 200 trials on five bimanual (two-armed) tasks. The thread calls that 3.9x higher.
This is the first head-to-head number against a named robotics model in what we have seen, after the earlier reports of Astra scoring on robot arm control.
Paxton says the baselines are soft
Chris Paxton's first reaction was blunt. "If you are a robot foundation model guy and not a deployment guy this must be really concerning," he wrote. Then he walked toward the baseline rather than the winner, saying most robotics baselines are not very good because they are contact light, and that a lot of pick and place tasks can be done well with SAM (Segment Anything, a vision model) plus a very light language model.
So the argument that followed the number was not about Astra at all. It was about whether five bimanual tasks that a segmentation model plus a small language model can mostly handle are a real test of anything. On that read, a general model winning here says more about the tasks than about the model.
What the numbers do not say
The thread text gives trial and task counts. It does not say whether the trials ran on real robots or in simulation, who picked the five tasks, or who ran them. Nobody in the sources has checked the numbers independently.
Separately, Yanjie Ze posted that GPT6 Astra solved a Rubik's Cube with robot hands. That is a video, not a benchmark, and the same caution applies.
Paxton made the caution explicit in a separate post, arguing that robots have to be out in the world doing a variety of things and cannot just learn in labs. His line on demos is worth holding onto while these Astra clips keep landing. When you see a demo, he wrote, remember it is the absolute best the robot has ever done, and if you did not see it do something, assume it cannot do that thing.
The interesting part of this week is not the 3.9x. It is that a benchmark win over a dedicated robotics VLA landed and the field's instinct was to go check whether the benchmark was hard enough.
GPT-6 Astra scores 46% vs 12% for MolmoAct2, a state-of-the-art robotics VLA, across 200 trials on five bimanual tasks. That's 3.9x higher.
most of the robotics baseline are not so good -- theyre very contact light. A lot of pick and place tasks you can do pretty well with segment anything (SAM) and a very light language model.
when you see a demo you have to remember its the absolute best the robot has ever done, and if you didnt see it do something, you must assume it cant do that thing

