Chelsea Finn asks what robotics' RLHF moment looks like
A new blog post argues reliability, not capability, is what stands between robot demos and robots people trust.
Chelsea Finn and Perry Dong posted a blog on what it will take for robots to be broadly useful in the real world. Their framing question is blunt. What will be the "RLHF" moment for robotics, and what will it take to get robotics to where LLMs are today and beyond.
Reliability is the one of the biggest open challenges in AI right now.
What will be the “RLHF” moment for robotics? What will it take to get robotics to where LLMs are today and beyond?
it will become more of a bottleneck as we want systems to act with more autonomy and more trust
RLHF, reinforcement learning from human feedback, is the training step that turned raw language models into systems people were willing to use every day. Finn and Dong say robotics has no equivalent yet, and their post lays out their thoughts on the state of RL for frontier robotics models and what is missing.
Reliability is the bottleneck
Finn's own summary is about reliability rather than raw capability. Current models work out okay if a person is reviewing the outputs, she writes, using code drafting as the example. That stops being fine when the system is expected to act on its own.
"it will become more of a bottleneck as we want systems to act with more autonomy and more trust," Finn wrote.
That is the part that matters for anyone watching humanoid videos. A model that is right most of the time is a useful assistant when a human checks the work. The same hit rate on a robot that is left alone in a house or a warehouse is a different product entirely, and the gap between those two situations is where most robot demos live.
Shown versus claimed
There is no new robot here, no benchmark number and no result. This is a position piece from two researchers about where the field is stuck, posted by Finn and by Perry Dong on the same day. The sources do not say what specific method they propose, only that the post covers the state of RL for frontier robotics models and what they think is missing.
Worth reading if you have been arguing about why autonomy keeps stalling after the demo.

