Robotics Weekly

Figure's 4 hours in 30 homes, and the 56% argument

Figure posted four straight hours of its humanoid working in 30 homes it had never seen, and the field spent the day arguing about the half of the time it fails. Meanwhile a controlled benchmark put GPT-6 Astra at 35% and Claude Fable 5.1 at 15% on real robot tasks that human teleoperators finished every time.

The big one

Figure posts 4 straight hours of Helix 2.5 working in 30 homes

Figure put out four hours of Helix 2.5 in 30 homes with no additional training on those houses. Adcock separately posted this week's data numbers from the Index app: 53 minutes of real-world human video uploaded every second, 2.4M uploads in a week, and 115,000 weekly active users. He calls that the flywheel behind Helix, and says no other form factor can learn from human video at the same scale. A summary by @itsolelehmann that Adcock quoted says the same robot finished the jobs 9% of the time without the human-video pretraining and 56% with it.

Brett Adcock
@adcock_brett
X
If you want to watch something really boring: here's 4 hours of our humanoid doing zero-shot work across 30 rental homes
Sep 18, 2026 · View on X
Brett Adcock
@adcock_brett
X
Our humanoid robots are learning from the largest installed dataset in the world: humans operating in the real world
Sep 18, 2026 · View on X

Read the full story

Shown this week

Benchmark: Astra 35%, Fable 15%, human teleoperators 100%

Pantograph ran GPT-6 Astra and Claude Fable 5.1 on eight manipulation tasks on its Pandroid robots. Astra completed 35% of attempts, Fable 15%, with p roughly 0.001, while human teleoperators (people driving the robot remotely) completed every task. @addadryanw said the gap to teleop is the real headline, and Paxton agreed that none of it means contact-rich manipulation is close to solved.

Pantograph
@pantographPBC
X
Astra completed 35% of attempts, Fable 15% (p≈0.001). Human teleoperators completed every task.
Sep 17, 2026 · View on X
Chris Paxton
@chris_j_paxton
X
time for some more soul searching on the part of my field i think
Sep 18, 2026 · View on X
Chris Paxton
@chris_j_paxton
X
It's fair that absolutely none of these results imply contact rich manipulation is anything close to solved
Sep 18, 2026 · View on X

Read the full story

Teleop and other arguments

Keerthana Gopalakrishnan calls Figure's pretraining comparison a weak baseline

Tony Zhao kicked it off by pointing at a 237/420 success rate he says he took from Figure's blog, and said failing half the time is not "doing real useful work." Goldberg asked the obvious follow-up: what are the failure modes. Gopalakrishnan, who had defended Figure for being transparent about its numbers, then said she found a devil in the details: the blog compares an Index-pretrained model against a from-scratch model, a baseline she says nobody really uses since VLAs (models that turn camera images and text into robot actions) were invented. Kalouche replied that Figure leads the way in overclaiming and gaming demos. Paxton split the difference, calling it great that Figure is showing real success rates while arguing that even a 1% failure rate in a home is too much. The real story: Zhao says his complaint is aimed at the marketing line in the video, not the engineering. Gopalakrishnan's point is the sharper one. She notes Figure is not claiming Index pretraining beats VLM pretraining, only that it beats no pretraining, which she says should be obvious. Nobody in the thread has answered Goldberg's failure-mode question yet.

Keerthana Gopalakrishnan
@keerthanpg
X
They are comparing index pretrained model vs a from-scratch model - which is a weak baseline that no one really uses after VLAs were invented?
Sep 17, 2026 · View on X
Simon Kalouche
@simonkalouche
X
You’re defending Figure? They lead the way in over claiming and gaming demos…..
Sep 18, 2026 · View on X
Chris Paxton
@chris_j_paxton
X
no one wants a $20,000 robot that breaks 1% of their dishes or gets stuck doing the laundry 1% of the time either
Sep 18, 2026 · View on X

Read the full story

Also on the timeline

Ken Goldberg's group pitches robots with zero human demonstrations

Animesh Garg summarized work from Ken Goldberg and collaborators as a third path: no human demonstration data, multi-agent coding harnesses, video-to-sim system identification, and offline agent self-correction that ships as deterministic edge code. His framing is that the AI writes the robot's code rather than running its inner loop.

Q-Planning takes credit card insertion from 25% to 80%

Varun Giridhar and Animesh Garg add a Q-function (a model that scores how good an action is) on top of a large policy like pi-0.5, then update it online from both successful and failed robot rollouts. On RoboPapers episode 105 they report a credit-card-into-wallet task going from 25% to 80% in a few iterations.

Watney raises $80M and claims four nines in data centers

Watney announced an $80M Series A, over $100M total, and says it has served the largest hyperscalers since 2025 with more than four nines of reliability across hundreds of thousands of hours in customer facilities. It claims the largest fleet of dexterous robots running 24/7 in the United States. Lukas Ziegler noted the company was folding hotel laundry two years ago.

Dreamscale Labs bets robot inference belongs in the cloud

A new YC company launched saying onboard compute is getting too expensive for frontier robot models, so it is building real-time inference infrastructure that runs the model off the robot while still meeting motion deadlines. Jiafei Duan said deploying at massive scale needs the cloud.

Get the next one by email

Robotics every day from the people building it. The demos, the deployments and the arguments worth your time. Every claim links back to the engineer.