Malik vs Isola: can an LLM actually control a robot?
Jitendra Malik turned the GPT-6 Astra robotics hype into a testable challenge and Phillip Isola took it. Plus Astra's real2sim demos, Lightwheel's 100,000 hours of first-person video, and a Boston Dynamics spinout for theme parks.
The big stuff
Malik challenges Isola: control a quadruped with an LLM. Isola will try
Malik's argument: the Astra robotics demos are simple pick and place with parallel jaw grippers (two-finger pinch grippers), which mainly shows that LLMs (large language models like GPT) can plan, while the hard part is high frequency control of torques, forces and contact. So he posted a challenge: prompt an LLM to output the high frequency control commands for a legged robot on varying terrain, pointing at RSS 2021 and CoRL 2022 work he calls five-year-old technology. Isola said he likes the challenge and will see if his team can try, and agreed that real dynamic and dexterous control has not been shown yet. Goldberg said locomotion can be learned but wagers that agents using control theory will make more progress on fast, reliable industrial manipulation this year than pure model-free methods (learning control by trial and error with no physics model).
The tasks that are demonstrated are simple pick and place tasks with parallel jaw grippers. LLMs can do planning, and the impressive demos in these tasks primarily show that.
I’d wager we’ll see more progress on fast and reliable manipulation this year using Agentic Robotics than with pure model-free methods.
People making grand claims about #robotics really need to spend time with actual robots. Debug the controller yourself.
Demos and launches
GPT-6 Astra reportedly built a sim from robot video and painted the Golden Gate
Guo says he gave GPT-6 only multi-view camera images, robot actions and one prompt, and it did real2sim (rebuilding a real scene as a physics simulation) end to end: calibrated the cameras, built object assets, ran physics system-ID (fitting the simulator's physics to the real data), ran MuJoCo (a physics simulator) and rendered the result in Blender. @cdngdev says he handed Astra a robot, a brush and a camera, asked it to paint the Golden Gate Bridge, and it worked out how to control the robot and improved over its attempts. Kuan constrained Astra to MediaPipe's output schema for 3D hand pose on egocentric (first-person) data and found high reasoning mode took 3 minutes per frame against 20 milliseconds for MediaPipe. @darpinian argued it will be hard for anyone without 400k GPUs to compete and that GPT-7 might be a fine-tune away from beating every robotics model.
It calibrated the cameras, built the object assets, performed physics system-ID, ran MuJoCo, and rendered in Blender.
High reasoning took 3 min per frame, while MediaPipe took only 20 ms per frame.
i played around with it on a couple personal projects and it's great and a real improvement but not the total step change people are claiming (or alternately skill issue)
Research
Lightwheel releases 100,000 hours of first-person video, Figure's Index hits 69,900 weekly users
Lightwheel has released 100,000 hours of egocentric video (first-person footage, like a head-mounted camera) covering human tasks from cleaning and tidying to assembly, construction and retail, and it is on HuggingFace now. RoboPapers' case for egocentric data: cheap to collect, covers the breadth of real human tasks, and true to real physics. Separately, Adcock says Index hit 69,900 weekly active users last week and is taking in 35 minutes of uploaded data every second, up from the 30 minutes per second he cited at launch alongside 16M uploads and $15M paid out.
it truly seems as if the end of the robotics data shortage is in sight
Last week we hit 69,900 weekly active users
Quick hits
Boston Dynamics spinout Dynamic Creatures emerges to build theme park character robots
Dynamic Creatures, co-founded by former Boston Dynamics chief strategy officer Marc Theermann and former RAI Institute research leader Farbod Farshidian, came out of stealth as Boston Dynamics' official entertainment and hospitality partner and says it is already working with a major theme park and a retailer on interactive character robots. The platform is called SnowJay, and operators lease characters instead of buying them upfront. Backers listed are Eniac Ventures, Kindred Ventures, Heliad Sunshine Lake and Bluegrass Ventures, no amount given. Note that the poster, Lukas Ziegler, says he is an angel investor alongside Marc Raibert.
Haoyang Weng open-sources HDMI, which learns whole-body robot skills from human video
Haoyang Weng rewrote and open-sourced HDMI, a framework for learning whole-body interaction skills for a robot directly from human videos with no manual reward engineering, now with an mjlab backend and support for multiple motions on a single object. The original paper reported 67 door traversals, 6 real-world tasks and 14 in simulation, and Weng says the released code replays the video with the README command.






