Developer says his Astra harness cut a robot task to under a minute
@calixo888 posted a cup-in-cup run he says dropped from about an hour to under 60 seconds after two days of side projects.
@calixo888 says a harness he built around GPT-6 Astra took robot task execution from about an hour to under a minute. He describes it as two days of side projecting with the model.
our harness reduced robot task execution from 1hr to
The task was small. The prompt he posted was "put the grey cup inside the orange cup. do it fast, make no mistakes". That is one pick and place with a clear goal state, not a long chore with steps that can go wrong in twenty ways.
What a harness is doing here
A harness is the code wrapped around a model that feeds it images, decides how often it gets to look and think, and turns its output into arm commands. Same model, different plumbing. When people report order of magnitude speedups on Astra runs, the plumbing is usually what changed, because the slow part of these demos has been the model stopping to reason between every small move.
That matches the pattern from earlier Astra tests. The most quoted result so far was the Ethernet cable run that took two hours and four human nudges. Speed, not capability, is the thing everyone is poking at.
Shown versus claimed
One task, one developer, his own numbers. There is no success rate across trials in the post, no comparison run someone else set up, and no detail in the source about what the harness actually changes. So the honest read is a single reported result on an easy task, not a benchmark.
Chris Paxton, responding in the same conversation, expects the latency problem to fade. "This will get a lot better," he wrote, pointing at OpenAI's new voice interfaces as evidence the company can handle dynamic, streaming, real-time data. That is a bet about the direction, not a measurement.
The interesting question is whether a harness that turns a one hour cup stack into a one minute cup stack survives contact with a task where the robot has to recover from its own mistakes. Nobody has shown that yet.
This will get a lot better.

