Figure's 30-home demo and the 237/420 fight
Figure teased an "AI breakthrough," then showed 4 hours of a humanoid working in 30 rented Bay Area homes with no new training. Tony Zhao pulled the number out of the blog post: 237 successes out of 420 tries.
Read the whole week in one goFigure's 237 out of 420 became the week's fight
The big one
Figure says Helix 2.5 worked in 30 unseen homes with no new training
Figure rented 30 homes in the Bay Area and ran its humanoid there with no additional training, posting 4 hours of zero-shot footage plus a detailed write-up. Adcock framed it around whether a humanoid can walk into a house it has never seen and get to work on its own, whole body, unassisted. Tony Zhao then pointed at a number in the blog: 237 successes out of 420 attempts. Adcock replied that Figure published per-task percentages for everything.
We rented 30 homes in the Bay Area and are doing tasks without any new training
Love my friends at figure, but failing half the time is not “doing real useful work.” See the 237/420 success rate below taken from the blog.
Let’s not shame people for being transparent about their success rates. In an industry where many people are over claiming and gaming demos, transparency and grounding claims in rigor should be applauded and incentivized.
Shown this week
Chris Paxton says Astra's run at the top is already over
Chris Paxton wrote "Well astra being on top didn't last long" and, in a follow up, said of whoever displaced it: "If you're wondering who these guys are, former Tsinghua/BIGAI team". Earlier the same day he replied to someone's results with "I am glad to see these negative results." The real story: these are reactions, and the posts they point at are not in front of us. We cannot say which benchmark changed hands, what the numbers were, or what the negative results were about. What is visible is the tempo. A model sits on top, a team most people cannot name knocks it off within days, and a senior researcher is thanking someone for publishing a failure. Negative results in robotics are still rare enough to be worth a thank you.
Well astra being on top didn't last long
I am glad to see these negative results.
If you're wondering who these guys are, former Tsinghua/BIGAI team
Teleop and other arguments
Lerrel Pinto: billion-dollar startups are demoing what kids with a few GPUs can do
Pinto did not name a company or a demo. Shafiullah replied to him quoting the line “The robot arrives with no additional training and starts doing useful work” along with a link. Nothing in the posts says which company or robot either of them means, so the target is unconfirmed. The real story: this is a swipe, not evidence. In a separate exchange, Tony Zhao replied to @OwenBrakes with "ya just want to be nice. you can just say things these days", and the post being discussed is not visible, so what he was defending or mocking is unknown. Treat all of it as field chatter until someone names the demo and shows the numbers.
Fascinating how multi-billion dollar startups can claim a major AI breakthrough, then demo what smart kids with a few GPUs can do today 🤦
ya just want to be nice. you can just say things these days
Also on the timeline
MessyMem gives a mobile robot memory that updates when it acts
Jeannette Bohg's group posted MessyMem, a 3D scene graph that a VLM updates from what each interaction reveals, plus keyframes for details the graph does not store. Reported numbers: 84% on cluttered pick and 99% on locked cabinets with all pieces, dropping to 58% and 43% in ablations, and 80% task progress over 25 tasks and 3+ hours.
Ken Goldberg says AI agents are tuning robot controllers in real package sorting
Ken Goldberg posted that "Agentic Robotics", where AI agents design and fine-tune robot control systems, is being used in the real world to improve e-commerce package sorting. The real story: that is the whole claim in the sources. No numbers, no video, no detail on which system or how much it improved, and no one in the sources says this is a first. Treat it as a pointer worth watching, not a result you can check yet.
Jiafei Duan: standardized robot evaluation is still one of the hardest problems
Jiafei Duan praised a robot evaluation effort for being run "in such a systematic and reproducible way," and used the moment to point at a bigger gap: there is still no settled way to compare robot policies fairly. The real story: the sources here do not show what was evaluated, which models were compared, or what the numbers were. All we have is a researcher in the field saying standardized evaluation remains unsolved. That is a real complaint, and it is why cross-model claims in robotics are so hard to check.
MIT's FloatForm robot boats latch into bridges on open water
Lukas Ziegler wrote up FloatForm from MIT CSAIL and the Senseable City Lab, published in Nature: 21 cm square boats that self-assemble into connected structures using an origami-inspired latch with permanent magnets, closing gaps of 10 to 15 cm. Holding a connection draws no power, and the assembled structure then navigates as one vessel.




