Turning internet videos into robot hand data
A retargeting algorithm called Do as I Do claims it can turn ordinary monocular video, even generated video, into robot hand motion that actually runs on hardware. Dexterity data has been the bottleneck, and this is a swing at it.
Shown this week
Do as I Do turns plain video into executable robot hand data
Bhawna Paliwal, Haritheja Etukuru, Will Liang and Mahi Shafiullah presented Do as I Do, an algorithm that reconstructs monocular RGB video (ordinary single-camera footage) and retargets it onto robot hands. RoboPapers says it beats the state of the art and works even on generated video.
Real-world dexterity data, including things like finger-pose estimates, is often slightly off, making it physically invalid and hard to execute on real hardware and hard to learn from.
FANUC shows a cobot you talk to, trained before the arm existed
At FANUC Europe's AMB booth, a CRX cobot took plain-language commands, with vision picking out the object and NVIDIA's GR00T vision-language-action model planning the grasp. Lukas Ziegler says the team built against Isaac Sim as a digital twin because the physical hardware was not available yet, then used that same simulation to prepare the fine-tuning data.
So the robot learned its task before the robot existed, and the environment that taught it also stood in for the machine itself.

