Researchers argue RL should replace IK for low-level control
A one-line question from Xiaolong Wang turned into a fight over whether model based control has a future at all.
Xiaolong Wang asked Chris Paxton a short question about a robot demo, and it turned into one of the week's sharper arguments about how robots should be controlled.
The question was plain. "What does RL add here compared to doing IK?" IK is inverse kinematics, the classic way of solving which joint angles put the hand where you want it. Paxton's answer was reactivity. He said RL deals better with many object masses, and with what happens on a sudden impact or a caught arm, and that whole body RL handles those cases more gracefully than IK without a lot of engineering effort.
Liao goes further
Qiayuan Liao took the position all the way. He said any online optimization for low-level control, naming IK, ID and MPC, should be abandoned. His reasoning is that an RL policy can fundamentally track more precisely, more robustly and more adaptively to unknown disturbances and uncharacterized noise, especially on low-cost hardware, and he says that holds even on a manipulator.
Paxton staked out similar ground elsewhere in the thread, writing that the future clearly is not in model based control, even for wheeled robots.
The future clearly is not in model based control, even for wheeled robots
Wang is not convinced
Wang replied to Liao with four words. "Not super sure about this." That is the whole pushback, but it is from the person who started the thread, and it leaves the strongest version of the claim unsettled.
Shown versus claimed
The honest part of the exchange is Paxton's own caveat. He noted that, as others pointed out, Dyna does not show any of these things, so it is less clear, and then said he still thinks it is the right strategy. In other words the argument for replacing IK with RL is being made from impact and disturbance cases that the Dyna footage in question does not actually demonstrate. It is a bet on where control is heading, not a measured result.
That matters because the two camps are not arguing about a benchmark number. Nobody in the thread posted tracking error, impact recovery rates or hardware costs. It is three researchers disagreeing about whether solvers that run optimization online at control time are worth keeping around once learned policies get good enough, and about whether cheap hardware is where learned control wins first. Liao thinks they are not worth keeping. Paxton thinks the direction is right even without the evidence. Wang is not sold.
Any online optimization (IK, ID, MPC) for low-level control should be abandoned.
As others point out, Dyna doesnt show any of these things, so its less clear. But I do think its the right strategy

