ARLI: Fixing the Missing State in Asynchronous Robot RL. Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency • Brian Zhu, Momen Khalil, E. Harrison, Emanuele Poggi, et al. • Siemens, UC Berkeley, Microsoft, ETH Zurich • arXiv preprint • 2026. The robot moves before the answer arrives. Imagine a robot aligning a connector with a socket. Its camera captures an image, its vision-language-action model starts inference, and its controller continues executing commands generated during the previous inference cycle. By the time the new prediction arrives, the robot is somewhere else. That is manageable if the learning system understands what happened during the intervening time. But suppose its reinforcement-learning policy sees only the original image. The same apparent observation can now precede different physical outcomes, depending on commands already sitting in the execution queue. The learner is being asked to explain transitions using an incomplete description of what caused them. ARLI—Asynchronous RL with Intermediate Information—addresses this mismatch by coupling a frozen VLA with a small learned steering policy that receives information unavailable when the large model started computing. Its central contribution is not simply keeping the robot moving. It is making that moving, delayed system more learnable. First, an attribution clarification: despite the supplied Stanford AI Lab attribution, the linked paper is a Siemens–UC Berkeley-led collaboration, with Microsoft and ETH Zurich also represented in the author affiliations. For this episode of Embodied AI 101, the important question is where ARLI intervenes. It does not make the foundation model instantaneous, and it does not demonstrate unrestricted end-to-end RL of arbitrary frontier policies. Instead, it changes what a lightweight adaptation policy knows, and when that policy makes its decision. That distinction is useful well beyond this particular implementation. When computation overlaps physical execution, the inference pipeline itself becomes part of the control problem. The missing state is often sitting in your command queue. Concurrent control predates VLAs. In Thinking While Moving, Ted Xiao and colleagues studied agents that select their next action while previous actions are still unfolding. Their formulation explicitly incorporates information about the preceding action and action-selection delay; their practical representations include a “vector-to-go” describing the remaining commanded movement. The underlying lesson is that the controller’s unfinished business can matter as much as the robot’s measured configuration. Consider a deliberately simple example. Two executions have exactly the same physical state when a camera image is captured. In one, the command queue contains a sequence that will move the gripper left. In the other, it contains a sequence that will move the gripper right. A slow policy receives the same image in both executions and produces the same next chunk. That chunk will nevertheless begin from different locations. Nothing mysterious has happened to the physical dynamics. The learner’s observation has omitted a causal variable: the commands that will execute before its new decision takes effect. This also explains why adding more training data need not solve the problem cleanly. More data can teach the critic an average over the hidden queues encountered during collection. But that average is not equivalent to conditioning on the actual queue. As the policy changes, the distribution of those queues can change too. My interpretation is that this creates two distinct difficulties. One is ordinary uncertainty about future physical events. The other is avoidable uncertainty about the controller’s own commitments. We should not ask a neural network to infer the second from pixels if the software already knows the answer. Earlier delayed-control work makes this distinction explicit. Delay-Aware Model-Based Reinforcement Learning for Continuous Control, by Baiming Chen and colleagues, constructs an augmented-state formulation for delayed actions. Delay does not make an MDP representation impossible; it changes what must be included in that representation. There is an important corollary here. Stopping the robot while inference runs is not automatically the mathematically correct solution. It changes the execution regime. A policy that was demonstrated with continuous motion may now encounter repeated deceleration, holding, and restarting. Conversely, asynchronous execution is not automatically a learning solution. It removes those pauses while leaving the learner responsible for commands selected earlier. The design problem is therefore more precise than “synchronous versus asynchronous.” We need to choose an execution schedule and an information state that describe the same physical process. Why the useful intervention point is the noise input. ARLI builds on Diffusion Steering via Reinforcement Learning, or DSRL. In Steering Your Diffusion Policy with Latent Space Reinforcement Learning, Andrew Wagenmaker and colleagues treat a pretrained diffusion or flow policy as a transformation from latent noise to robot actions. Instead of sampling that noise from its usual Gaussian distribution, an RL-trained policy chooses it. The generator’s weights remain unchanged; adaptation occurs through the distribution of its inputs. A useful mental model is a learned action interface. The large policy already knows how to turn a latent input and an observation into a coordinated movement. The small policy learns which latent inputs produce useful outcomes. This is different from asking RL to discover manipulation directly in joint-command space, and different from adding a correction after action generation. It is also computationally attractive: DSRL can learn without differentiating through the pretrained generator. The original method requires black-box access to its noise-to-action interface rather than access to its weights or every denoising operation. For asynchronous inference, that interface has another advantage. The steering decision need not always be made when visual-language processing begins. If noise is consumed only by a later action-generation stage, there is an opportunity to delay the small policy’s decision without delaying the overall output. That is the opening ARLI exploits. The implementation uses DSRL-SAC and repeated latent actions. In the corresponding DSRL construction, one per-step latent vector is repeated across the chunk’s noise tensor, reducing the dimensionality of the RL action space. This does not mean repeating the same physical command: the generator still produces a temporally structured action sequence. There is a subtle trade-off worth keeping in mind. A learned generator can make exploration more structured, but its noise interface is not an arbitrary control interface. Some corrections may be easy to express through that latent space; others may be inaccessible or difficult to discover. Nor does freezing the generator constitute a safety guarantee. Changing its input distribution can change its behavior substantially. The relevant engineering question is whether the available steering directions cover the corrections the task needs—not merely whether the large model’s parameters remain fixed. ARLI’s two additions: the committed actions and a later observation. The method’s core information state contains three ingredients: the observation that launched VLA inference, the previously committed actions spanning the inference window, and a more recent observation obtained before the action expert needs its steering noise. Think of these as answering three different questions. The first observation tells the steering policy what situation the large model is processing. The committed actions describe what the robot is already scheduled to do. The later observation provides feedback about what actually happened while computation proceeded. Those roles are complementary. A queue predicts commanded motion but cannot report an unexpected slip. A fresh image reports the visible scene but need not explain all commands that remain committed. And retaining the original observation preserves context for a generator whose expensive processing began earlier. The scheduling insight is to run the lightweight policy as late as possible while still supplying its noise before action generation requires it. The authors’ presentation emphasizes that the shorter inference time of the RL policy allows it to use a newer state than the large model initially received. Here is how I would reason about that schedule when implementing it. Start with the time at which the next chunk must become executable. Work backward through the action expert, the steering network, and any required data movement. That determines the latest useful observation deadline. The slow backbone can start earlier, while the old command queue continues running. The objective is not to make every component consume the newest possible image. That would simply restart expensive computation. The objective is to deliver new information to a component that can still affect the upcoming output. The real-robot configuration uses 60 Hz control, 50-action predictions, 20 executed actions, ten steps of total delay, and seven steps between the steering observation and execution. The latter two correspond to approximately 167 and 117 milliseconds: the newer observation saves 50 milliseconds, not the entire VLA delay. That distinction matters. The remaining seven-step interval is not a claim that the small steering network itself takes seven control steps to run. It is the remaining decision-to-execution budget, including downstream action generation. For a practitioner, I would therefore distinguish three measurements: total inference latency, the age of the steering observation when execution starts, and the interval before another steering decision can affect the robot. Optimizing one does not automatically optimize all three. As a hypothetical example, accelerating the backbone may allow the whole call to start later. Accelerating the action expert may additionally allow the steering policy to observe later. Those changes can have different consequences even if they save the same number of milliseconds. ARLI makes that distinction operational: information freshness depends on where computation sits in the pipeline. Why smooth chunk boundaries are not enough. The closest deployment-side companion is Real-Time Chunking, or RTC. In Real-Time Execution of Action Chunking Flow Policies, Kevin Black, Manuel Galliker, and Sergey Levine frame overlapping action generation as an inpainting problem. Actions guaranteed to execute during inference constrain the beginning of the new prediction, while guidance encourages consistency in the remaining overlap. The method operates at inference time and does not require retraining the policy. RTC and latency-aware RL address different questions. RTC asks whether the new trajectory is compatible with the trajectory already being executed. An RL learner asks which decision will improve future reward given the situation in which that decision takes effect. My reading is that the distinction resembles the difference between enforcing a boundary condition and exposing a state variable. A trajectory generator can obey a committed prefix without the critic explicitly knowing which prefix was committed. Conversely, a critic can receive the queue while the generator still produces an awkward handoff. That is why I would resist treating smooth rollout videos as evidence that the learning state is sufficient. Smoothness is valuable, but it does not establish that the actor and critic possess the information needed for credit assignment. There is also a computational cost to consider. Training-Time Action Conditioning for Efficient Real-Time Chunking shows an alternative: train the policy to condition directly on action prefixes, avoiding the vector-Jacobian computations required by inference-time inpainting guidance. That approach changes the training recipe rather than merely wrapping a frozen generator. A natural follow-up experiment would combine late steering with a prefix-conditioned action expert. The hypothesis would be that cheaper continuity enforcement leaves a smaller blind interval after the steering observation. But that is a proposed combination, not an ARLI result. It would need to be evaluated for both latency and steering expressiveness. A prefix-conditioned generator could alter how effectively latent noise controls the remaining trajectory. The broader lesson is to measure continuity, information freshness, and adaptation capability separately. Calling the entire stack “real-time” conceals distinctions that matter for learning. “Near-Markov” is not the same as “equivalent to zero delay”. The user’s summary describes ARLI as restoring the Markov property. The more careful interpretation is that it addresses information omitted by naive asynchronous conditioning; the project itself uses the language of near-Markovian structure. An augmented state can be exactly Markov under an appropriate delayed-system formulation. But a pair of images, joint measurements, and queued commands is not automatically a complete physical state. Occluded contacts, object deformation, unmeasured forces, and uncertain timing can still matter. It is useful to separate two questions:. Does the representation support learning a consistent transition model? And how much performance is lost because information arrives too late to influence a decision?. Those are not the same question. A perfectly specified stochastic decision process can still impose a substantial penalty for delayed observation. The theoretical background also involves how action-chunk data are collected. In Decoupled Q-Chunking, Qiyang Li, Seohong Park, and Sergey Levine discuss open-loop consistency: conditioning a dataset on an entire action sequence should agree with the transition distribution obtained by actually executing that sequence open-loop. If later actions were chosen in response to intermediate events, conditioning on those actions can select particular events and bias that comparison. For intuition, imagine a reactive controller that turns left only after observing an obstacle. Selecting every recorded sequence containing a left turn also selects episodes in which that obstacle was observed. That is different from deciding in advance to turn left regardless of what appears. ARLI’s guarantee requires open-loop-consistent data covering optimal chunks and optimal completions of committed prefixes. Its performance loss is bounded through a delayed-oracle gap accumulated over the discounted horizon. I read this as an information-and-coverage result, not a blanket convergence guarantee for practical neural SAC. A final thought experiment makes the limit concrete. Suppose a latch randomly changes which insertion direction is valid immediately after the steering image is captured. No policy restricted to that image can react to the realized change before another observation becomes actionable. Better bookkeeping removes avoidable uncertainty. It does not reveal events that have not yet been observed. What the experiments establish—and what they do not. The simulation suite includes Kinetix’s , , and , together with AlohaTransferCube. The project compares ARLI with naive asynchronous DSRL and DSRL combined with RTC. Simulation averages three seeds, with 50 evaluation rollouts; its learning curves use environment steps rather than a wall-clock training budget. For interpreting the study, I would view these environments as complementary probes rather than one undifferentiated benchmark score. A dynamic low-dimensional task can expose timing failures clearly. A visual manipulation task tests whether the mechanism remains useful behind a larger perception-and-action stack. Neither alone establishes broad generality. Together, they provide a more informative test of the central hypothesis: does supplying the learner with intermediate information help under asynchronous execution?. The real experiments use a bimanual UR5e cell and task-adapted π₀.₅ policies. The tasks are connector assembly, placing a shoe into a narrow bag, and repositioning a bag with both arms. The reported improvement is roughly 40% to near 100% success after 100–125 training episodes, measured using a ten-episode rolling average. That is encouraging evidence of online improvement. It is not an industrial reliability estimate. A rolling training curve answers whether recent experience is becoming more successful. For deployment, I would additionally want a frozen-checkpoint evaluation over independently sampled initial conditions, repeated training runs, and longer operation. Those measurements answer different questions and should not be substituted for one another. The learning budget also needs its initialization context: the real-world recipe includes 1,200 initial chunk-level transitions and 24,000 offline updates. My takeaway is not that this invalidates the sample-efficiency result. It is that “one hundred episodes to improve” should not be interpreted as the complete cost of producing the system. A replication should account separately for task adaptation, initial experience collection, offline optimization, and subsequent online interaction. The actions-only ablation is competitive on assembly and shoe packing but weaker on bag placement. That is particularly informative because it discourages a one-size-fits-all explanation. My hypothesis would be that some tasks are dominated by predictable queued motion, while others benefit more from observing how the scene actually evolved. The result motivates that hypothesis; it does not by itself identify the physical cause. The throughput metric also deserves careful reading. It is estimated from the last 20 episodes, assuming failures take twice the average successful duration. I would treat that as a useful policy-level comparison, not a complete work-cell productivity measurement. A production assessment should include reset labor, recovery, verification, intervention, and downtime. A faster successful rollout and a more productive autonomous station are related but distinct accomplishments. Finally, the authors describe cases matching or exceeding standard RL in idealized no-latency settings. That compares trained systems, not information-theoretic optima. A true zero-delay oracle could imitate a delayed policy if doing so were advantageous. Beating a particular no-delay learning baseline does not imply that withholding information is inherently beneficial. The strongest interpretation remains narrower and more useful: latency-aware information design can make online improvement work where naive deployment-and-learning combinations struggle. What I would build first. For reproduction, the public starting points are DSRL’s π₀ implementation, Physical Intelligence’s , and the RTC Kinetix repository. The DSRL repository provides a JAX implementation with Aloha and LIBERO examples; the RTC repository provides simulation code and pretrained assets. As of September 28, 2026, I did not find an ARLI-specific training repository linked from the project page. Those upstream repositories should therefore be treated as building blocks, not a verified one-command reproduction. The following is my proposed implementation sequence, rather than additional reported experimental results. First, instrument the execution timeline. I would log sensor capture time, inference submission, backbone completion, steering invocation, chunk completion, and actual controller execution. Host-side function duration is not enough. The important quantity is the age of information when its consequences reach the robot. I would also log the queue as a timestamped object, not merely as an array attached to a training sample. A command’s meaning depends on when it is scheduled and whether it was actually executed. Second, make queue snapshots atomic. The steering observation and committed-action sequence must refer to a consistent execution schedule. If one thread advances the queue while another constructs the RL input, the learner can receive a plausible-looking but causally incorrect state. This is the kind of bug that can survive ordinary unit tests. I would create deterministic replay tests in which the same sensor stream and command schedule must reconstruct exactly the same augmented observation. Third, give the critic the relevant temporal information. It would be a mistake to make the actor latency-aware while leaving the critic to average over hidden command queues. The value function should describe the decision process the actor actually operates in. For a first implementation, I would favor explicit timestamps, masks, and queue features over asking a large recurrent encoder to discover the scheduling semantics implicitly. A more expressive representation can come later, after the timing is demonstrably correct. Fourth, define replay transitions around decisions, not convenient software callbacks. The reward interval, selected latent action, execution offset, terminal flag, and next augmented observation need a coherent interpretation. This becomes especially important when a task ends while a chunk remains queued. I would explicitly test reset boundaries. Commands produced before a reset must not leak into the next episode, and observations captured after a reset must not be attached to the preceding decision. Fifth, separate learning compute from control deadlines. Online updates can compete with inference for memory bandwidth and accelerator time. A policy that meets its deadline when the learner is idle may behave differently during sustained training. My evaluation would therefore include the full training system under load. I would report deadline misses, observation age, and queue underruns alongside success. Finally, I would run controlled ablations with the same scheduler and training budget: old observation only, queue augmentation only, later observation only, and both. I would then repeat that comparison with continuity enforcement enabled and disabled. The purpose is to establish which information fixes which failure—not merely to reproduce a favorable final curve. Where I would expect the design to run out of room. The first boundary is architectural access. DSRL’s noise interface can be exposed by a relatively simple black-box API. Late steering demands more: the implementation must allow the steering decision to arrive before the relevant action-generation stage consumes it. A service that accepts one observation and returns one finished chunk may hide precisely the scheduling boundary we need. That is an integration issue, not necessarily a learning issue. Before adapting the algorithm, I would inspect the inference interface and verify that the intended intervention point actually exists. A second boundary is the remaining blind interval. If disturbances require correction faster than the action expert can finish, a later noise decision may still arrive too late. Residual RL offers a different intervention. Residual Off-Policy RL for Finetuning Behavior Cloning Policies learns corrections around a behavioral-cloning policy’s actions, rather than steering the generator’s input noise. My architectural interpretation is that these approaches trade different forms of authority. Latent steering can influence a coherent generated trajectory. A downstream residual controller can potentially modify an already available command using newer feedback. Which is preferable depends on whether the task needs better trajectory selection, faster local correction, or both. A combined system is an interesting research direction, but it raises additional questions. Which component receives credit? Can the residual layer undo assumptions made by the chunk generator? What happens when the two adaptations compensate for each other rather than improving the underlying behavior?. A third boundary is variable latency. I would not assume that training with one fixed delay establishes robustness to jitter. Queue length, observation age, and remaining compute can vary independently. Testing only average latency can miss the behavior that occurs during rare long calls. My most revealing stress test would inject disturbances at several phases: before initial observation, during backbone computation, after the steering observation, and during execution. That would map the system’s actual reactive window. Finally, adaptation cannot replace ordinary control protections. I would retain independent command limits, deadline handling, and stop behavior. A useful learned correction mechanism should operate inside a well-defined execution envelope, not be asked to discover that envelope through failures. The architectural takeaway. ARLI’s most useful contribution is a question:. What information becomes available while the large model is computing, and which component can still use it?. That question leads to a different way of designing robot policies. Instead of treating inference as one indivisible function call, we can examine its internal deadlines. Some decisions must happen early. Others can wait. Information that arrives too late for one component may still be valuable to another. For RL, the companion lesson is that the command queue is not merely implementation detail. It can be part of the state needed to explain outcomes. My conclusion is therefore more specific than “latency is solved” or “frontier VLA post-training is unblocked.” ARLI presents a compelling mechanism for improving asynchronously executed diffusion- and flow-based robot policies: make prior commitments visible, and postpone the lightweight adaptive decision until fresher information can influence it. The robot still cannot know the future. But the learner should at least know what the robot is already committed to doing—and should not decide earlier than necessary.