Robotics Weekly

Astra hits 79% on RoboMME at 3.63 calls per episode

A three-tier stack puts a frontier model in the robot loop only when a small monitor says it is needed.

Bingao Chen reports 79.13% success on RoboMME with GPT-6 Astra, using an average of 3.63 Astra calls per episode. The run covers 800 official test episodes across 16 tasks.

Bingao Chen
@bingao_chen
X
let Astra decide what to do, and let the small model decide when to ask it again
Sep 21, 2026 · View on X
Bingao Chen
@bingao_chen
X
RoboMME challenges robots to act on information beyond the current camera view
Sep 21, 2026 · View on X

RoboMME is built around tasks where the answer is not in the current camera frame. Remember where a hidden object went. Follow a sequence that was demonstrated in a video. Count how many times an action has already been completed. That rules out a policy that just reacts to what it sees right now.

How the stack is put together

The setup has three tiers. Astra does the reasoning. A small visual monitor watches the scene. A VLA (a model that turns camera images and text into robot actions) does the moving. Chen describes the trick as letting Astra pick the plan and letting the small model decide when Astra needs to be asked again.

That is what the 3.63 number is for. Calling a frontier model on every timestep is the obvious way to get a smart robot and the obvious way to make it unaffordable. Under four calls for a whole episode is a claim about cost as much as capability, and cost is what decides whether a model like Astra can sit inside a control loop at all rather than being a demo.

What has not been shown

This is a thread, not a paper, and there is no independent replication in the sources. Chen says the evaluation used the official test split, which at least makes the 79.13% comparable to other RoboMME numbers, but nobody outside has run it. The thread also does not say which VLA or which small monitor model was used.

Response so far is light. Yinpei Dai called it interesting work applying GPT-6 Astra to RoboMME. It lands in a busy few weeks of outside groups poking at Astra, including a earlier benchmark that put Astra at 35% against human teleoperators at 100%. Different benchmark, different task set, so the two scores are not comparable.

Get the next one by email

Robotics every day from the people building it. The demos, the deployments and the arguments worth your time. Every claim links back to the engineer.