Robotics Weekly

RPG robot watches one demo, practices in sim, goes 30 for 30

A real-to-sim-to-real pipeline that rebuilds the task in simulation, practices it, and updates a shared skill set from its own failures.

Yen-Jen Wang has posted RPG, a pipeline that takes a real-world demonstration, rebuilds the task in simulation, practices it there, and then runs the improved skill back on hardware. Wang calls it real-to-sim-to-real, and the loop is spelled out as demo, reconstruct, guided practice, self-improve, go real.

Yen-Jen Wang
@wangyenjen
X
A truly agentic robot shouldn’t just know what to do. It should also figure out how to get better.
Oct 7, 2026 · View on X

The pitch is that a robot should not only copy what it saw. It should work out what it is bad at and drill that part. RPG reconstructs practice tasks from the demonstration, uses the demo video itself to steer how the skill improves, and keeps updating a shared skill system from its failures rather than throwing each attempt away.

The numbers

Two results, and they are very different kinds of result. On real robots, RPG hits 30 of 30 successes across 3 tasks. In simulation, it reaches 95.0% success across 22 manipulation tasks after 15 rounds of self-improvement.

Shown versus claimed matters here. The big headline figure, 95% across 22 tasks, is simulation only. The hardware number is perfect, but it covers three tasks and thirty attempts, which is a small sample for anyone trying to judge whether this generalises. Wang's post does not say which tasks they were or what hardware ran them.

Why the loop is the interesting part

The idea of training in sim and deploying on a real robot is not new. What RPG adds is the direction of travel. Instead of a human building a simulated task and hoping it transfers, a single real demonstration defines the task, the simulator rebuilds it, and the robot generates its own practice inside that reconstruction. The demo video stays in the loop as guidance rather than being consumed once.

The shared skill system is the other piece worth watching. Failures feed back into a common pool rather than a single task policy, which is the part that would make this scale past three tasks. Nobody has shown that yet. The announcement is a thread, and the sources here do not include a paper, code or a public benchmark run.

For now it is a clean result with an honest gap between the sim column and the real column. The next thing to look for is whether the 22-task number survives contact with hardware.

Get the next one by email

Robotics every day from the people building it. The demos, the deployments and the arguments worth your time. Every claim links back to the engineer.