RL vs IK debate narrows to whose model is cruder
Qiayuan Liao, Roei Herzig and Kevin Zakka traded arguments about whether a learned policy or a classical controller is the one making bad assumptions.
The thread about whether reinforcement learning should replace inverse kinematics (solving the joint angles that put a robot's hand where you want it) has moved on from "which one wins" to "which one is lying to you about the robot". It started with the earlier exchange between Chris Paxton, Qiayuan Liao and others, and the latest round is three people arguing about model fidelity.
Liao says solve it offline
Qiayuan Liao's position is that real robot performance is not decided by numerical accuracy. He lists what breaks a clean controller in practice, friction, inertia mismatch, state-full friction and sensor noise, and argues that the usual fixes, disturbance observers and online adaptive controllers, all rest on strong assumptions and will not scale. His conclusion is blunt. If you can simulate the thing, solve it offline, humanoid or tabletop 6DOF arm.
Herzig says that trade costs precision
Roei Herzig pushed back. RL trains in simulation, so it can only be as accurate as that simulator. Randomizing the sim, the standard trick for making a policy survive the real world, buys robustness and pays for it in precision. For precise tracking on a fixed-base 6DoF arm, Herzig says IK and MPC (model predictive control, which plans actions a short way ahead using a model of the robot) still win.
Zakka turns the argument around
Kevin Zakka's reply is the sharpest line in the thread. IK assumes the robot rigidly goes where you command it, and stiction, backlash and flex all break that assumption. MPC, he says, is also only as good as the model it runs on, and that model is typically cruder than an RL simulator because it has to run in real time. So "only as accurate as the simulator" cuts both ways.
That is the whole shape of the disagreement. Everyone agrees the controller is bounded by the model behind it. They disagree about whether the richer, slower model inside a physics sim or the fast, simplified model inside a real time controller is closer to the actual machine.
Shown versus claimed matters here. Nobody in this exchange posted a benchmark, a tracking error figure or a video. There is no experiment attached, just three practitioners with different priors about where the error comes from on hardware they each know well. Treat it as a useful framing of the question, not a settled answer.
IK assumes you rigidly go where commanded. Stiction, backlash and flex violate that.
Randomizing the sim makes the policy robust, but costs precision.
If we can effectively simulate it, why don’t we solve it offline


