Robotics Weekly

Reality Check launches with 14,400 real robot rollouts

A public leaderboard for robot manipulation, run on real hardware rather than in simulation.

Nicolas Keller has launched Reality Check, a leaderboard and what he calls the first public robot manipulation benchmark. It is built from 14,400 real-world rollouts across four models.

Nicolas Keller
@Nicolas_Keller
X
The Era of Evals is coming to AI robotics.
Sep 29, 2026 · View on X

The setup covers multiple tasks and data regimes, varied object placements, and reports confidence intervals. The runs were done on FR3 Duo stations. "The Era of Evals is coming to AI robotics," Keller wrote in the launch post.

Real robots, not simulation

The detail that matters here is the word real. These are physical rollouts on hardware, not simulated episodes, and each result comes with a confidence interval rather than a single headline number. That combination is what has been missing while the field argued for months about which demo video to believe.

It also means the numbers are expensive to produce. 14,400 rollouts on real arms is a lot of robot hours, which is part of why nobody had published a public manipulation leaderboard like this before.

MolmoAct2 on top

Jiafei Duan, who has been vocal about how robot demos get presented, says MolmoAct2 came out as the state of the art on the new benchmark. "This is the year of RoboEval!" he wrote.

Duan has been pushing the same point from the other direction, posting uncut real-time footage of robot rollouts to show how slow the real thing looks compared to a sped-up clip.

Shown versus claimed

The sources do not include the task list, the four model names beyond MolmoAct2, or the scores themselves. Keller's post announces the benchmark and the scale of the evaluation, and the thread continues past the opening post. What a reader can take from this today is that the leaderboard exists, it is public, and it was run on real FR3 Duo stations with 14,400 rollouts behind it.

The open question is whether other labs submit to it. A public benchmark only settles arguments if the people being argued about agree to be measured by it, and nobody knows yet whether that happens.

Jiafei Duan
@DJiafei
X
Exciting to see MolmoAct2 being the SOTA for this new public benchmark!
Sep 29, 2026 · View on X

Get the next one by email

Robotics every day from the people building it. The demos, the deployments and the arguments worth your time. Every claim links back to the engineer.