Astra drove a robot arm from one human video, Xiao says
A first-pass success on one task, a self-written real2sim pipeline, and a benchmark post that says the gap is still ten times.
Wenli Xiao says his team dropped a recording of a human doing a novel task into the Codex app, prompted it to drive a robot arm the same way, and it worked on the first pass. He calls it physical ICL (in-context learning, where a model picks up a task from an example in the prompt with no retraining). That is one task, one report, and no success rate or number of takes appears in any source.
GPT-6 Astra just does physical ICL out of the box.
The real2sim piece
@huxiao93612565 posted a second result the same day. He gave GPT-6 Astra one video of a human hand, and says the model wrote an entire real2sim pipeline (turning a real video into a physics simulation) by itself, covering hand tracking, IK retargeting (mapping the human hand motion onto the robot hand) to a 44-DOF dexterous hand, and grasp refinement. Zero code from him, and he says he is open-sourcing everything. Animesh Garg replied that the real2sim progress is amazing and that it required so much effort as recently as last month.
Ken Goldberg framed the broader shift. He says agentic robotics provides the missing link between model-based and model-free methods, and that to deal with latency it is now generating executable ROS2 code (a common robot software framework). His line is that the frontier is now harness design, not the model.
The counterweight
Markus Wulfmeier posted updated robot deployment benchmarks for both Google DeepMind and OpenAI models. Astra is finally overtaking Gemini with a nearly 8% jump since Sol, and it has overtaken Gemini Robotics ER 2. But the Gemini model in the comparison is only the flash model, and Wulfmeier flags the price difference, $10 and $50 per million tokens in and out for GPT versus $0.75 and $3.75 for Gemini flash. He says a large gap remains to his team's internal models, at ten times lower error rates.
Keerthana Gopalakrishnan replied asking for a link and for the Gemini Robotics ER versus Gemini flash result. No link or breakdown appears in the sources.
So the shown versus claimed split is clean. What is shown is one first-pass arm demo, one self-written simulation pipeline being open-sourced, and a benchmark where a general model is climbing but still trails specialist internal models by an order of magnitude. That is consistent with the earlier reports of Astra's jump on robot arm control. Whether any of this holds across many tasks, nobody has shown yet.
To address latency AR is now generating executable ROS2 code.
A large gap still remains to our internal models (x10 lower error rates) but I'm excited to see the recent advances.


