ENPIRE: Can Coding Agents Improve Real-World Robot Manipulation?. ENPIRE: Agentic Robot Policy Self-Improvement in the Real World • Wenli Xiao, Jia Xie, Tonghe Zhang, Haotian Lin et al. • NVIDIA, Carnegie Mellon University, UC Berkeley • CoRL 2026 / arXiv • 2026. Can a coding agent improve a real robot’s manipulation policy without a roboticist rewriting every experiment?. ENPIRE’s answer is yes—after people help establish a reliable experimental environment. Its contribution is a framework for automating policy development through physical trials, not a robot foundation model that independently acquires general dexterity. And the headline 99% success rate needs its qualifier immediately: the project reports pass@8, allowing up to eight feedback-conditioned retries per subtask within a rollout. That is a recovery-aware completion metric, not 99% first-attempt reliability. For practitioners, the interesting question is therefore not whether the demonstration looks autonomous. It is whether the surrounding infrastructure makes autonomous experimentation useful, trustworthy, and economical. This episode of Embodied AI 101 examines the September 20, 2026 paper revision, the public implementation, and related work on code-based control and real-world reinforcement learning. The distinction we will keep returning to is simple: automating research inside an engineered robot cell is a meaningful achievement, but it is not the same as eliminating the engineering of that cell. What the Paper Did. ENPIRE organizes robot-policy development into four modules. The environment module provides resets and outcome verification. Policy improvement changes a controller or its training procedure. Rollouts execute candidate policies and collect evidence. Evolution compares experiments and carries useful changes forward. The public implementation supports both Python-based control programs and neural-policy training through an actor–learner pipeline. The important architectural decision is a division of authority. A policy researcher can change how the robot solves the task, but should not change what counts as solving it. ENPIRE’s published research contract freezes the environment, reset procedure, verifier, safety limits, and evaluation seeds. It asks the agent to record a falsifiable hypothesis, make a policy-side change, test it, inspect evidence, and retain it only if it passes the configured acceptance gates. Failed trials must remain in the record. That boundary is more consequential than the acronym. Imagine a pin-insertion experiment in which an agent notices repeated failures near contact. It might propose a different controller, change a training objective, or collect more useful trajectories. But allowing it to lower the insertion-depth threshold would destroy the comparison. The research system needs freedom to improve the policy and restrictions that preserve the meaning of improvement. The physical demonstrations include pushing a T-shaped object into alignment, pin insertion, GPU insertion, zip-tie fastening, and cutting a zip tie. The reset examples are themselves substantial manipulation programs: recovering objects, positioning them for another attempt, and checking readiness. The zip-tie verifier combines evidence from two camera views rather than accepting a potentially misleading overlap in one image. Notice the decomposition. A reusable reset can place an object near the start of the difficult contact phase. That lets policy search concentrate its physical trials where precision matters. My interpretation is that this is sensible task engineering, not a shortcut that invalidates the work. But it changes what we should credit to the learned component: success of the assembled system does not mean every stage was learned by one policy. There is also a practical distinction between the coding agent and the controller it develops. ENPIRE’s repository describes Python programs as a primary policy representation while retaining a separate optional learning subsystem. The language-model agent is doing software development and experimental orchestration; it need not be the component choosing every low-level motor command. For a working roboticist, that opens several possible adaptation paths. You could keep an existing controller and automate its parameter experiments. You could expose a behavior-cloning pipeline and let the agent investigate data selection. Or you could allow changes to an RL learner while holding the physical interface fixed. These are different experiments with different risks; “agentic self-improvement” should not collapse them into one capability. The clearest reported hardware scaling result uses bimanual YAM stations. Moving from one to eight agent–robot workers reduced Push-T convergence from approximately five hours to two, and near-perfect pin-insertion performance from more than 90 minutes to roughly 40. The pin-insertion research objective includes achieving 50 consecutive successes. Those results establish that the loop can produce useful controllers and that parallel resources can shorten research time. They do not, by themselves, tell us whether an agent discovered a better learning algorithm, whether collaboration was essential, or whether the resulting controller generalizes outside the tested cell. We need separate evidence for those stronger interpretations. Finally, “minimal human effort” needs an operational definition. In his behind-the-scenes explanation, coauthor Jim Fan explicitly acknowledges the substantial preparation and the continuing presence of human operators watching the robots. Autonomous policy iteration should therefore be distinguished from an entirely unattended facility. What the 99% Result Actually Measures. A recovery-aware metric is not inherently suspicious. If a robot can detect a bad grasp, reposition, and complete an assembly without assistance, that recovery is part of its competence. The problem arises when a completion metric is interpreted as a precision metric—or when neither includes the cost of recovery. Consider two hypothetical controllers. One normally succeeds immediately. The other frequently uses most of its retry allowance. They could have the same final completion percentage while differing dramatically in cycle time, contact count, component wear, and opportunities for damage. A production decision cannot treat those systems as equivalent. For that reason, I would want three curves rather than one headline number: first-attempt success, completion versus retry budget, and completion versus elapsed time. The last curve is especially useful because “one retry” may be a tiny corrective movement in one system and a lengthy regrasping sequence in another. There is also no legitimate shortcut from ENPIRE’s pass@8 figure to a first-attempt probability. The project explicitly describes retries conditioned on earlier failures, not independent samples. Applying an independent-trials formula would substitute a different experimental model for the one being reported. The 50-success research target raises a separate statistical issue. Here is an illustrative calculation, not a confidence interval for ENPIRE’s headline result: even if a frozen policy passed 50 independent, predeclared tests, the exact one-sided 95% lower confidence bound on its success probability would be about 94.2%. That is encouraging evidence, but it would not establish 99% reliability with high confidence. If the successful window is selected while the policy keeps changing, the simple fixed-policy calculation is not applicable in the first place. This distinction matters because an experiment can legitimately use a rolling success window as a stopping rule. What it should not do is quietly treat that stopping rule as a separate validation study. I would freeze the selected policy, start a fresh evaluation, and specify the trial count before looking at results. The verifier deserves equally serious treatment. A frozen reward prevents the policy researcher from directly rewriting the scoreboard, but freezing does not make a detector infallible. A classifier can remain systematically wrong under an unfamiliar occlusion, lighting condition, or failure geometry. My proposed audit would deliberately collect ambiguous near-successes: a pin that looks aligned but is not seated, a strap that overlaps the head without passing through it, or an insertion that reaches the expected pose but is mechanically unacceptable. Those examples test the boundary that matters most, rather than average classification accuracy over easy frames. I would also report reset reliability separately and include it in end-to-end accounting. As a hypothetical illustration, a reset that succeeds 98% of the time followed by a task policy that succeeds 99% of the time yields only about 97% successful complete cycles before other failure modes enter. That is not an estimate of ENPIRE’s performance. It is why the denominator matters. A deployment metric should begin when the system accepts responsibility for the next job, not only after the difficult preparation has already succeeded. How Much of the Improvement Comes From the Agent?. ENPIRE sits at the intersection of two established capabilities: learning manipulation through real-world feedback, and using language models to write robot-control code. Neither capability began here. HIL-SERL, for example, demonstrated precise manipulation through reinforcement learning supported by demonstrations and human corrections. Its reported evaluations include near-perfect success, cycle-time improvements, and practical training durations—typically one to two-and-a-half hours on most tested tasks. That provides important context: rapid, high-success real-world RL was already possible with an engineered system and human assistance. ENPIRE’s distinctive question is closer to this: how much of the experiment-design and algorithm-adjustment work can move from the human researcher into an agent-controlled loop?. That is a valuable question even if the final controller resembles something an experienced engineer could have produced. Automation does not have to invent a new objective function to be useful. Repeatedly finding a competent configuration with less expert attention could be enough. The related Probe, Learn, Distill work, or PLD, makes another distinction clear. PLD trains residual specialists around a frozen VLA, collects deployment-relevant recovery trajectories, and distills those trajectories back into the generalist. It explicitly studies a route from specialist experience to improved model weights. That gives us a useful way to interpret the user-facing foundation-model narrative. An automated physical research system could become a source of valuable training data or specialist recipes. But generating stronger task-specific behavior and improving a broadly generalizing foundation model are separate experimental claims. I would want a subsequent distillation study, with held-out tasks, before treating the former as proof of the latter. On the code-generation side, CaP-X already investigates how abstraction level, execution feedback, and visual grounding affect coding agents for manipulation. It compares agents using controlled primitive interfaces and includes human-written reference solutions. Its results show why the available tools are part of the method, not incidental plumbing. This changes how we should construct a fair ENPIRE baseline. A weak baseline would give an agent one prompt and no useful diagnostics, then compare it with an agent allowed repeated experiments, better tooling, and extensive evidence. A stronger comparison would hold the robot interface and total budget fixed while changing the research strategy. I would include an experienced human using the same tools, a conventional parameter-search procedure, and independent coding agents without shared discoveries. Each answers a different question: labor substitution, search efficiency, and collaboration benefit. The paper’s exploratory trace contains a useful example: adding behavior-cloning regularization coincides with a 10.8-percentage-point improvement, whereas later batch-size and controller adjustments contribute around one percentage point each. Those are observed search milestones, not automatically isolated causal ablations. A plausible explanation for the larger gain is that imitation regularization discourages a policy from abandoning useful demonstrated behavior while pursuing sparse rewards. But that explanation should be tested by rerunning the change from matched checkpoints and data, not inferred solely from its position in a successful development history. The small gains need even more caution. Near saturation, a percentage point can be valuable, but it can also be difficult to distinguish from evaluation variation. I would repeat complete research runs as well as final-policy evaluations. The unit of variability is not only the robot episode; it is also the agent’s sequence of proposed experiments. One revealing ablation concerns observation access. On simplified physical Push-T, native vision reaches success first, but text-only research beats a configuration that obtains visual information through a separate callable module. The practical inference is not that images are unnecessary. It is that evidence delivery has a cost. A tool that supplies useful information too slowly or in an awkward format can make the overall research process worse. Before upgrading a backbone, I would inspect whether failures are observable, timestamps are aligned, and the agent can identify which policy produced which trajectory. Robot Fleets Buy Time, Not Necessarily Efficiency. The fleet result is promising, but “faster” needs a resource denominator. Using the reported approximate Push-T timings, one station for five hours represents five allocated station-hours. Eight stations for two hours represent sixteen. That is roughly 3.2 times as much allocated hardware time for a shorter wall-clock wait. This calculation concerns reserved capacity, not measured active execution time. The distinction matters because the project also reports declining per-robot utilization as the fleet grows, alongside higher token consumption. Agents spend time examining logs, developing code, and coordinating rather than continuously executing experiments. Whether that trade is attractive depends on your bottleneck. A lab with otherwise idle machines may happily spend more aggregate capacity to obtain a result before morning. A team constrained by hardware availability, fragile parts, or consumables may prefer slower but more economical search. This is why I like the decision to instrument resource use. Robot utilization, token consumption, and time to a successful policy expose different bottlenecks. They should not be collapsed into one notion of intelligence or efficiency. Fan’s technical explanation similarly emphasizes physical execution, compute activity, and token expenditure as distinct resources. However, a one-, four-, and eight-worker experiment is not enough to establish a universal scaling law. Nor does a speedup establish that communication between workers caused it. An especially informative control would run eight independent research workers and select the first successful result. Compare that with eight workers allowed to exchange findings, using the same total budgets and initial conditions. The difference would help isolate the value of collaboration from the value of parallel search. I would also separate two operational questions. Is the next experiment ready when the robot becomes free? And is the next experiment worth running?. A scheduler can improve the first without improving the second. Keeping expensive hardware busy with weak hypotheses is not necessarily progress. For deployment, I would track useful validated improvements per robot-hour alongside utilization. The engineering opportunity suggested by these results is therefore broader than adding more agents. Better experiment queues, smaller diagnostic artifacts, asynchronous analysis, and disciplined branch selection might recover throughput without expanding the fleet. That is a hypothesis for system builders to test, not a demonstrated ENPIRE result. Generalization: Policies, Recipes, and Tool Stacks. There are several different things that might generalize in this system. A policy might handle new initial conditions. A reset program might recover from a broader set of failures. A research recipe might help another task. Or the entire framework might port to a different robot. Those are not interchangeable. ENPIRE reports transferring written research experience from pin insertion to GPU insertion. The transfer setup retains an explicit summary while removing raw prior trajectories, logs, and checkpoints. This is evidence for transferring a research recipe through context, not direct transfer of the same trained motor policy. That is still interesting. A useful record of which objectives stabilized learning or which diagnostics exposed failures could save substantial engineering effort. But I would evaluate it against carefully chosen controls: generic advice, a shuffled summary, an expert-written checklist, and no summary. Otherwise, it is difficult to tell whether the benefit comes from specific transferable experience or simply from giving the new agent more guidance. The simulation evidence broadens the task setting. Figure 6 shows approximately three-quarter success for ENPIRE, compared with roughly half for GR00T N1.5 and roughly one-quarter for the depicted CaP-X variant on the selected RoboCasa tasks. These are substantial descriptive differences, not marginal one-point gains. The reported evaluation uses matched initializations and the same success predicate, with 40 episodes per task. That controls an important source of variance. But a system-level comparison should remain a system-level conclusion. Comparing generated programs that can orchestrate perception and planning with a learned-policy execution pathway does not isolate the value of agentic research alone. For that, I would additionally equalize the available tools and compare development strategies around the same starting controller. Nor should success in selected kitchen simulations be read as broad physical deployment validation. I would want new objects, independently varied contact conditions, shifted cameras, and another station before asserting transfer beyond the original operating envelope. A closely related later project, ASPIRE, explores persistent libraries of reusable robot skills and explicitly separates debugging seeds from held-out evaluation. Its project page also acknowledges that real-world deployment still requires reliable success detection, resets, monitoring, and calibration. It is useful adjacent evidence about skill reuse, not an independent replication of ENPIRE’s hardware results. The common research opportunity is clear: preserve what the system learns without overfitting the next task to yesterday’s debugging session. For practitioners, the artifact worth transferring may be a tested diagnostic or a robust recovery routine, not necessarily a bigger policy checkpoint. Can Another Lab Reproduce It?. As of October 7, 2026, ENPIRE has a public NVIDIA repository under an Apache-2.0 license. It includes environment examples, orchestration interfaces, and policy-learning components. That is substantially more useful than a video-only release. But available code does not mean complete reproduction of every demonstrated task. The real-world workflow documentation explicitly says that GPU and zip-tie code-as-policy scripts from internal deployments are not included. It also says the pin-insertion learner and actor are complete while users must provide the robot-side bridge; relevant robot-side launchers for GPU insertion and zip-tie experiments are likewise absent. A physical Push-T environment and supervisor are included. Those omissions matter because the physical feedback interface is central to the paper’s contribution. Reconstructing a bridge, reset, or verifier is not merely changing a file path. It can alter episode boundaries, timing, failure recovery, and the data distribution seen by the learner. There are nevertheless good reproducibility decisions in the release. The PLD actor–learner runtime has an isolated dependency environment, and the repository documents implementation provenance and rules for preserving existing behavior during refactoring. These choices reduce some of the ambiguity introduced by packaging a research system after the experiments. Hyperparameter reproducibility requires more than exposing defaults, however. For an adaptive research process, the relevant artifact is the configuration selected for each reported run, together with its data history and code revision. The published research contract asks for hypotheses, diffs, source commits, resolved configurations, checkpoint references, budgets, and decisions—the right kind of record for this problem. I would distinguish that logging requirement from evidence that every historical experiment has been released as a replayable package. A competent team needs both the mechanism and the records. There is also unavoidable physical setup. The calibration guide requires a correctly scaled printed board and manual placement. That is a useful reminder that a software clone cannot reproduce a station’s geometry. My reproducibility assessment is therefore mixed but constructive. Another capable robotics group has enough public material to understand and adapt the approach. Reproducing the exact end-to-end headline across all tasks would require additional implementation and experiment-specific artifacts. For a serious replication, I would preserve the agent’s model version, prompts, tool definitions, initial repository state, and complete sequence of accepted and rejected changes. Reproducing the final policy alone would test the controller; reproducing the development process would test ENPIRE’s central claim. What This Means for Practitioners. I would consider ENPIRE first for a repetitive laboratory or workcell task where the scene can be restored reliably, failures are observable, and physical experiments are affordable. I would not begin by promising an overnight general-purpose robot researcher. I would begin with one narrow task and ask whether automated experimentation can reduce the amount of expert attention needed to improve it. The first deliverable should be a trustworthy evaluation loop. Before giving an agent permission to modify training code, run the reset and verifier repeatedly without policy search. Deliberately create bad states. Check whether failures are identified, whether the station can recover, and whether the recorded result agrees with an independent assessment. Then make the policy interface explicit. Decide which changes are allowed: parameters, data sampling, network training, controller logic, or tool composition. Each additional degree of freedom makes the search more expressive and the attribution problem harder. I would keep hardware supervision separate from the experimental agent. The repository itself requires motion authorization, station and calibration checks, and stopping on unsafe or invalid states. Those controls should be enforced by the deployment system rather than treated as aspirations in a prompt. For the initial comparison, use the same environment and resource budget for a conventional baseline and the agent-managed procedure. Count expert intervention minutes as carefully as robot minutes. A system that needs fewer algorithm edits but substantially more recovery assistance may still be useful, but it is solving a different labor problem. I would also score the whole operating cycle. Include failed resets, timeouts, rejected parts, recovery time, and consumable replacement. If the application involves irreversible operations, ask how new material enters the loop. A policy’s success percentage is only one component of throughput. Once a promising controller appears, stop development and validate it separately. Use held-out initial conditions, run on another day, and, if possible, move to another calibrated station. Keep the distinction between development feedback and acceptance testing visible in both code and reporting. For teams building foundation models, the most attractive possibility may be an automated producer of specialist experience: difficult contacts, near-failures, and successful recoveries. PLD provides a related example of studying how specialist-generated trajectories can improve a generalist. Whether ENPIRE can make that pipeline broadly scalable remains a question for additional experiments, not something established by its task-completion headline. My overall assessment is positive, with a specific scope. ENPIRE makes a credible case for treating policy-development code as something an agent can improve through real physical feedback. The evidence is less decisive about collaboration efficiency, broad generalization, and deployment-grade reliability. The strongest takeaway is not “the robot no longer needs an engineer.” It is that some engineering work becomes amenable to automated search once the experiment is well specified. For a roboticist choosing what to reuse, I would prioritize the immutable evaluation boundary, auditable experiment records, and recoverable physical interface. Those are the foundations that let better agents become useful later. ENPIRE’s most important product is not the 99% number. It is a concrete demonstration of what must surround a robot before autonomous policy research becomes a meaningful experiment.