Berkeley and Siemens release ARLI, RL for laggy robot models
E Harrison posted the method, with Sergey Levine and Andrew Wagenmaker both amplifying it the same day.
A team from Berkeley and Siemens released ARLI, short for Asynchronous RL with Intermediate Information, a way to reinforcement-learn a VLA (a model that turns camera images and text into robot actions) when the model is too slow to react to what the camera sees right now.
The problem is delay. Running VLA inference asynchronously, meaning the next chunk of actions gets computed while the current one is still playing out, cuts the delay a robot feels. E Harrison says the cost is that it breaks the Markovian assumption, the idea in reinforcement learning that the current observation contains everything the policy needs to decide. If the action you are about to take was computed from an image that is already out of date, that assumption is gone, and standard RL fine-tuning stops being well posed.
A small policy steering a big one
Sergey Levine describes the fix as a two-model setup. A large robot foundation model does the heavy lifting, and a small RL policy runs alongside it. Because the small one is fast, it gets to look at more recent images, and it nudges the big model toward better behavior. Levine frames the question as how to run RL with real-time chunking, the technique of generating actions in overlapping blocks so the robot never stalls waiting on a forward pass.
Andrew Wagenmaker, who also amplified the work, puts the same issue plainly. To get smooth execution, actions have to be computed from past observations, so the question is how much that delay hurts RL fine-tuning and whether you can still do it.
Levine says the collaboration with Siemens was led by Brian Zhu, Momen Khalil and Emanuele Poggi from Siemens, with E Harrison from Berkeley.
Shown versus claimed
The posts describe the method and the setup, not a headline number. None of them state a success rate, a task list or a robot platform, so how much ARLI actually buys you on hardware is in the paper rather than the thread. Treat this as plumbing, not a new capability.
That plumbing matters anyway. Latency is one of the quieter reasons a big model that looks strong in evaluation looks worse when it is driving a real arm, because the world moves during the gap between seeing and acting. Most published work on VLA fine-tuning quietly assumes that gap away. This one starts from the gap and works forward.
we figured out how to use a small RL policy with a large robot foundation model
to achieve smooth execution, actions must be computed from past observations

