SGS does gear meshing and nut threading with zero real data
A small change to what the robot practices in simulation, and the policies transfer to contact-rich assembly without a single real demonstration.
Octi Zhang has released Success-Guided Sampling, or SGS, a change to how a simulated robot picks what to practice. The resulting policies mesh gears, thread nuts and insert pegs zero-shot from RGB cameras, meaning no real-world data and no fine-tuning on the real robot. The same recipe also drives quadruped locomotion across terrains harder than before.
The list of things it does not use is the point. No demonstrations, no retargeting, no tactile sensors, no per-task reward tuning. Zhang describes the training as just PPO at scale, PPO being a standard reinforcement learning algorithm. His framing is that the missing piece was never a better RL algorithm, it was a small tweak to what the robot practices.
No demonstrations, no retargeting, no tactile sensors, no crazy reward tuning.
What the tweak actually is
Abhishek Gupta, who worked on the project, gave the clearest explanation. In reset-driven RL the robot gets dropped into a starting state over and over, and most of those states are either so easy it always succeeds or so hard it always fails. Either way the batch teaches it nothing. SGS adaptively concentrates the reset sampling around states with moderate empirical success rate, so the learning signal in each batch stays high. That is what lets it crack high-precision problems like contact-rich assembly.
Gupta also notes he was the skeptic on the project and the students pushed through anyway. He calls the method dead simple.
Shown versus claimed
Zhang says every video on the project site is at 1x speed. That matters more than it sounds. Contact-rich assembly is exactly the category of demo that usually arrives quietly sped up, because threading a nut slowly is not a good clip. Zhang is inviting people to check the policies themselves.
Mateo Guaman Castro, a co-author, posted his own thread framing it against the standard objection that precise manipulation needs real-world data.
The work is credited to @octi_zhang, @mateoguaman, @patrickhyin, @rosario_scaling and @iggydagnino, with Gupta. There is a paper and a project website. What nobody has shown yet is how the success rates hold up outside the clips, since the threads lead with behaviors rather than numbers.
The missing piece wasn't a better RL algorithm, but a small tweak to what the robot practices.
Sound like you were at the demo
One email a day. Under three minutes. The demos, the deployments and the part the press release left out.

