Robotics Weekly

Researchers argue RL should replace IK for low-level control

A one-line question from Xiaolong Wang turned into a fight over whether model based control has a future at all.

Xiaolong Wang asked Chris Paxton a short question about a robot demo, and it turned into one of the week's sharper arguments about how robots should be controlled.

The question was plain. "What does RL add here compared to doing IK?" IK is inverse kinematics, the classic way of solving which joint angles put the hand where you want it. Paxton's answer was reactivity. He said RL deals better with many object masses, and with what happens on a sudden impact or a caught arm, and that whole body RL handles those cases more gracefully than IK without a lot of engineering effort.

Liao goes further

Qiayuan Liao took the position all the way. He said any online optimization for low-level control, naming IK, ID and MPC, should be abandoned. His reasoning is that an RL policy can fundamentally track more precisely, more robustly and more adaptively to unknown disturbances and uncharacterized noise, especially on low-cost hardware, and he says that holds even on a manipulator.

Paxton staked out similar ground elsewhere in the thread, writing that the future clearly is not in model based control, even for wheeled robots.

Chris Paxton
@chris_j_paxton
X
The future clearly is not in model based control, even for wheeled robots
Oct 1, 2026 · View on X

Wang is not convinced

Wang replied to Liao with four words. "Not super sure about this." That is the whole pushback, but it is from the person who started the thread, and it leaves the strongest version of the claim unsettled.

Shown versus claimed

The honest part of the exchange is Paxton's own caveat. He noted that, as others pointed out, Dyna does not show any of these things, so it is less clear, and then said he still thinks it is the right strategy. In other words the argument for replacing IK with RL is being made from impact and disturbance cases that the Dyna footage in question does not actually demonstrate. It is a bet on where control is heading, not a measured result.

That matters because the two camps are not arguing about a benchmark number. Nobody in the thread posted tracking error, impact recovery rates or hardware costs. It is three researchers disagreeing about whether solvers that run optimization online at control time are worth keeping around once learned policies get good enough, and about whether cheap hardware is where learned control wins first. Liao thinks they are not worth keeping. Paxton thinks the direction is right even without the evidence. Wang is not sold.

Qiayuan Liao
@qiayuanliao
X
Any online optimization (IK, ID, MPC) for low-level control should be abandoned.
Oct 1, 2026 · View on X
Chris Paxton
@chris_j_paxton
X
As others point out, Dyna doesnt show any of these things, so its less clear. But I do think its the right strategy
Sep 30, 2026 · View on X

Get the next one by email

Robotics every day from the people building it. The demos, the deployments and the arguments worth your time. Every claim links back to the engineer.