Do as I Do turns plain video into robot hand data
A retargeting algorithm aimed at the part of dexterity data that usually breaks, physical validity.
A new algorithm called Do as I Do reconstructs monocular RGB video (ordinary single-camera footage, no depth sensor, no multi-camera rig) and retargets it onto robot hands. RoboPapers says it outperforms the state of the art and works even when the input video was generated rather than filmed. The authors, @bhawna_paliwal_, @HarithejaE, @willjhliang and @notmahi, walked through it on episode 107 of RoboPapers with hosts @chris_j_paxton and @DJiafei.
Real-world dexterity data, including things like finger-pose estimates, is often slightly off, making it physically invalid and hard to execute on real hardware and hard to learn from.
Why the data breaks
The pitch is not more data, it is data a robot can actually run. RoboPapers frames dexterity as the next frontier for robotics and says the data for it is subtly hard to scale. Finger-pose estimates pulled from real-world footage tend to be slightly off, and slightly off is enough to make a trajectory physically invalid, so it will not execute on hardware and a policy learns little from it.
That is a different failure than the usual complaint about robot data. It is not that the clips are scarce, it is that the hand poses look right to a human eye and are wrong by the standards of a high-degree-of-freedom hand trying to close a grasp. Do as I Do is aimed squarely at that gap.
Internet video as a source
@chris_j_paxton describes it as a retargeting algorithm for turning internet or generated videos into high-quality, executable robot data for high-dof hands. The generated-video part is the eyebrow raiser. If the pipeline holds up on synthetic footage, the supply of training material stops being limited by what someone filmed with a real hand.
Shown versus claimed. What is public so far is the podcast conversation and the authors' own account of how the method compares to prior work. The sources give no benchmark numbers, no list of tasks, and no independent replication, so the beats-state-of-the-art line is the team's claim relayed by RoboPapers rather than an outside result. Anyone weighing it should wait for the numbers.
Do as I Do is a retargeting algorithm for turning internet or generated videos into high-quality, executable robot data for high-dof hands.

