Physical Coding: Robots That Check Their Work—and Keep the Fixes. Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence • Hongcheng Gao, Jingjing Zhou, and colleagues • HexaFuture • arXiv technical report • 2026. The failure hidden inside a successful motion. Imagine asking a robot to place two components into a shipping box and close the lid. It picks up one component, deposits it, closes the box, and stops. The grasp worked. The placement worked. The closing motion worked. The task failed. Where should we fix that failure?. We could collect more demonstrations of the complete task. We could improve instruction conditioning. But we could also ask a more surgical question: where was the condition that should have prevented the robot from closing the box?. That question is the useful entry point into Physical Coding. The proposal connects explicit world state, executable procedures, verification, and persistent experience. HexaAnything is the authors’ agent implementation; its surrounding harness coordinates models and tools rather than treating an action policy as the entire robot intelligence stack. The formal paper, whose title appears above, was submitted on September 28, 2026. “Physical Coding” names the approach, rather than being the report’s full title. My reading is that this work poses a productive systems question: can a robot’s experience leave behind something more targeted than another demonstration? Could it produce a corrected condition, a better tool, a reusable recovery procedure, or training data that explains which decisions actually worked?. For this episode, we will separate three issues that are easy to blur together: what the execution architecture makes possible, what the experiments demonstrate, and what a convincing sim-to-real learning loop would still need to establish. The distinction matters. A robot recovering during one episode is not necessarily improving across episodes. A tool repository improving is not necessarily a model learning. And a program that passes software checks is not necessarily a program whose physical assumptions are true. Two programs, connected by evidence. Physical Coding divides the interface into Code as World and Code as Policy. The former holds task-relevant entities, relationships, constraints, and progress. The latter organizes actions, tool selection, checks, and recovery. Validated execution records can subsequently support memory, program changes, or model training. The report says the world representation is constructed from available observations rather than privileged simulator state; benchmark success is nevertheless assigned by the benchmark’s official evaluator. Those are separate information channels. To make this concrete, let us design an illustrative version of the shipping-box task. This is an implementation sketch, not a reconstruction of the authors’ exact software. The world record would contain identities for the two requested components, the box, and its lid. For each component, it would track the most recently supported location and whether the evidence is still usable. The task specification would distinguish “both requested components are inside” from the weaker condition “two objects appear near the box.”. The policy would not simply issue two placements followed by a closing command. It would first determine which components remain outside. It would select an available manipulation tool for one of them, execute that tool, and inspect the resulting evidence before advancing. Suppose the gripper releases a component over the box, but the next image is occluded by the wrist. There are now several different possibilities: the component landed correctly, it bounced out, or it remains caught on the gripper. Treating all three possibilities as “placement completed” would erase precisely the distinction the workflow needs. Instead, our illustrative policy would request a new observation. If that observation establishes the component’s location, execution can continue. If not, the task remains unresolved. HexaAnything’s specified verdicts distinguish successful checks, failed checks, inadequate evidence, blocked execution, and safety stops. That distinction suggests an important implementation discipline: unknown should not silently become true. I would also make evidence expire. A component observed inside the box before the robot bumped the table should not automatically remain certified afterward. The relevant question is not just whether a predicate was once supported, but whether intervening actions could have invalidated its supporting observation. This is where the representation becomes more useful than a prose checklist. A checklist can say “component placed.” An executable record can specify which component, which observation supported that conclusion, which actions occurred afterward, and which subsequent operations depend on it. The low-level tool can still be learned. We need not replace a capable manipulation policy with hand-written kinematics. In our example, the executable workflow delegates the motion while retaining responsibility for whether the task is progressing. The architectural attraction is therefore selective explicitness: make the decisions that need inspection accessible, without demanding that every contact interaction become symbolic. The relevant lineage—and a fair novelty claim. The first useful comparison is Code as Policies, by Jacky Liang and colleagues. That work already generated robot programs that combined perception calls, control interfaces, numerical computation, and control flow. It included feedback-oriented and reactive policies, not merely fixed action lists. Physical Coding should therefore not be introduced as the discovery that robots can execute language-model-generated programs. Statler, by Takuma Yoneda and colleagues, supplies another important precedent. It maintains an explicit world-state estimate and conditions subsequent actions on that estimate. Its reader–writer organization makes persistent state an active part of embodied reasoning rather than leaving the planner to reconstruct everything from an expanding interaction history. Voyager addresses a different part of the story. Its Minecraft agent accumulates executable skills and revises programs using execution feedback, without updating the underlying language model’s parameters. It is a clear example of an agent improving through persistent software artifacts rather than through weight updates. A particularly close empirical neighbor is CaP-X. It evaluates coding agents under different abstraction levels and feedback conditions. Its central finding is that human-designed abstractions substantially affect performance; multi-turn interaction, structured feedback, visual differencing, skill synthesis, and additional reasoning can recover some of the lost capability. CaP-X also studies reinforcement learning over the code-based interface. That comparison creates a demanding question for any new harness: how much intelligence lives in the model, and how much has been supplied through the tools?. An interface exposing “assemble the package” presents a very different problem from an interface exposing camera observations, grasp proposals, and bounded motions. Both systems might generate code. Their division of labor is not remotely equivalent. There is also a closely timed approach, RACaP, which separates code evolution from deployment. It evolves reusable policy APIs and associated infrastructure, then lets a reasoning-and-acting agent call frozen APIs during execution. Its design reinforces the importance of distinguishing a development-time programming agent from a runtime robot agent. Taken together, these precedents suggest a better novelty test than asking whether Physical Coding uses programs, memory, or feedback. Those ingredients already exist. The stronger question is whether their integration produces a more reliable and more learnable interface: one that exposes useful failure boundaries, preserves the evidence behind decisions, and supports improvements that survive beyond the rollout in which they were discovered. I would judge that integration by controlled comparisons, not by the breadth of the terminology. If a new representation helps, we should be able to identify what becomes easier to verify, what becomes easier to repair, and which costs move elsewhere in the system. What the manipulation experiments establish. RoboCasa365 is a simulation framework for household mobile manipulation, spanning atomic skills and composite activities. Its standard multi-task evaluation uses a 50-task target set; the name does not mean every reported experiment evaluates all 365 tasks. Here, the conditions share XR-1 as the action model and a maximum environment-step budget, with 50 seeds per task. Both coding-agent conditions use the reported GPT-5.6-Sol planner. The reported success rates are:. | Execution system | Atomic-Seen | Composite-Seen | Composite-Unseen | Overall | |---|---:|---:|---:|---:| | Native XR-1 | 78.0% | 54.8% | 34.3% | 56.6% | | General-purpose Codex agent | 81.7% | 59.8% | 34.1% | 59.5% | | HexaAnything | 80.9% | 61.5% | 38.3% | 61.1% |. These are the report’s task-level evaluation results, with overall success pooled across trials. The four-percentage-point improvement on unseen composites is encouraging, but it is not a deployment-level reliability result. The most useful baseline here is arguably the general-purpose coding agent. It asks a more specific question than the native-policy comparison: does a robotics-oriented execution harness add value beyond simply giving a capable coding model access to the action tool?. For my own reproduction, that would be the beginning of the investigation, not its end. I would first separate task decomposition from state maintenance. Give one condition the same sequence of local instructions but no persistent world record. If performance remains unchanged, the representation may be less important than the decomposition. Then I would separate verification from additional observation. A system that looks again after every action might improve even without sophisticated predicates. The relevant comparison would preserve the observation schedule while changing how observations are interpreted and retained. Next comes recovery. Does the system benefit because it detects a bad transition earlier, because it chooses a better repair, or because it simply receives more opportunities to try? These possibilities call for different measurements. Matching environment steps helps, but I would additionally measure wall-clock time, model calls, observation requests, and intervention opportunities. A planner might spend substantial computation deciding not to execute a physical action. That computation can be worthwhile; it should still appear in the accounting. Finally, I would want paired outcomes rather than only aggregate percentages. For each seed, which baseline failures became successes? Which baseline successes became failures? Were improvements concentrated in a few task families?. Those questions are especially important when the proposed advantage is compositional reasoning. A single overall rate can conceal both a useful recovery mechanism and a new failure mode caused by overcomplicated planning. The practical interpretation is not “code wins.” It is: this experiment provides a reason to investigate the execution layer while keeping the motor policy fixed. What actually evolves?. The model-update experiment produces HexaModel v0.1 by fine-tuning Qwen3.8-27B on a mixed corpus. Harness-derived data include about 9,300 RoboCasa365 judgment examples and 1,010 agent traces, contributing 38.8% of 178.6 million training tokens. In the same harness, overall success rises from 60.5% to 61.7%, and Composite-Unseen success from 37.3% to 39.5%. This is the right place to distinguish several possible explanations. A fine-tuned planner might become better at selecting tools. It might become more consistent about formatting calls. It might learn when another observation is useful. It might improve its task-state judgments. Or the general-domain portion of training might contribute capabilities that help all of those decisions. The authors’ project discussion explicitly notes that the mixed training recipe does not isolate the contribution of harness-derived data. The ablation I would want is therefore not simply “before training versus after training.” I would compare matched token budgets containing general data alone, state-judgment data, execution traces, and their combination. I would also test whether the trained model needs less assistance. Can it solve held-out compositions with fewer retries? Can it operate with a shorter history? Does it request observations more selectively? Does it retain the improvement when tool descriptions change slightly but their semantics remain fixed?. Those tests would help distinguish learning the workflow from learning the surface form of its transcripts. The tool-evolution evidence is different. In the cloth-folding study, a programming agent revises a reusable manipulation tool, including pinching and synchronized two-arm transport. A human selects image coordinates. After development involving privileged garment keypoints, a restricted reevaluation achieves four successes on five held-out seeds. That is an interesting assisted software-improvement experiment. It should not be mistaken for autonomous perception learning. For a robotics team, however, assisted tool improvement can still be valuable. Consider a hypothetical placement primitive whose approach trajectory works well for rigid objects but catches a deformable edge. A general revision to its approach and release behavior could improve many future calls without modifying the planner’s weights. The key evaluation question would be whether the change is reusable. Does it help new objects and placements? Does it preserve performance on the rigid-object cases that previously worked? How much operator effort was required to discover and validate it?. I would reserve separate labels for four outcomes: recovery within an episode, reusable memory across episodes, revised executable tools, and improved model parameters. Calling all four “self-evolution” may be convenient, but recording them separately makes the research much easier to interpret. The strongest future result would connect them: a failure produces a tested tool revision, that revision generates better traces, and training on those traces improves a planner on genuinely held-out tasks. A robot scientist needs more than successful grasps. The work also introduces PhyBench, a simulated laboratory covering spring stiffness, gravitational acceleration, and coupled-oscillator frequencies. With the planner identified as Opus 5.5, reported mean relative errors are 1.3%, 1.0%, and 1.7%, respectively, with ten valid runs per task. These errors are averaged over valid runs, not all attempted runs indiscriminately. Why is this a useful direction?. Return to our distinction between motion completion and task completion. In a manipulation task, we might ultimately care about object arrangement. In an experiment, the final deliverable is a numerical conclusion whose evidence must survive inspection. Here is how I would design a small spring-measurement exercise around that requirement. I would ask the agent to record an unloaded reference, choose several loading conditions, and retain each measurement together with the condition that produced it. Before accepting a reading, the workflow would need evidence that the apparatus was ready and that the measurement belonged to the current load rather than the previous one. The analysis program would operate on those recorded pairs. If the fit looked suspicious, I would want the agent to identify which observation should be repeated—not merely announce low confidence in the final answer. This example makes provenance operational. A suspicious data point should lead back to an image, an instrument reading, an action, or an unresolved state transition. Otherwise, the agent has a table of numbers but no practical way to investigate its own conclusion. I would evaluate such a system along two axes. One is whether it completed a valid physical measurement procedure. The other is whether the resulting estimate met the accuracy requirement. Keeping those axes separate prevents a misleading summary. A system can produce accurate estimates on the few runs it completes while failing to obtain usable evidence on most attempts. Conversely, it can execute a valid procedure consistently but need better measurement or analysis. The research opportunity is not simply to make a robot repeat a textbook experiment. It is to connect experimental decisions, physical execution, measurement records, and analytical code closely enough that a bad result can be diagnosed rather than discarded. What “sim-to-real” means here. It helps to contrast this approach with DrEureka. DrEureka uses language models to construct rewards and domain-randomization settings for policies trained in simulation and deployed physically. Its transfer mechanism directly concerns how simulation training prepares a motor policy for real-world conditions. The Physical Coding project reports real deployment on an AgileX dual-arm platform, with five of seven tabletop tasks succeeding in all three trials. The other two tasks—tossing blocks and folding clothes—succeed once in three trials each. Published timing references use different hardware and operating constraints. Its physical execution arrangement freezes validated tools between episodes and lets the runtime agent compose them. One reported tool revision replaces timer-based toss release with release based on measured arm position after observing firmware lag. That last example captures the kind of adaptation I find most interesting: an observed physical discrepancy motivates a localized change to an executable mechanism. But I would not interpret interface portability as equivalent to eliminating the dynamics gap. Suppose our shipping-box workflow runs in simulation and on a real robot. The same high-level procedure could remain useful while nearly everything beneath a grasp call changes: perception, calibration, trajectory generation, gripper behavior, and completion sensing. In that situation, what transferred was the organization of the task. Whether the motion capability transferred would require a different experiment. The same distinction applies to predicates. A simulator adapter and a physical adapter might both expose “component inside box,” yet implement that condition differently. One could use exact geometry; the other could use a partially occluded camera view. Identical function names would not establish equivalent evidence. For an interface-level transfer claim, I would therefore test the observation and action contracts themselves. Do units and coordinate frames agree? Does “completed” mean the controller stopped, or that the requested physical condition holds? Can both backends represent uncertainty rather than forcing every query into a Boolean?. For a policy-level transfer claim, I would freeze the relevant policy and specify exactly which adapters, calibrations, and physical data were allowed to change. For an adaptation claim, I would report the physical interaction budget and the human work involved. A method can be valuable even if it needs engineering intervention, but the intervention belongs in the result. The real-robot sample size also deserves perspective. As a simple calculation, a genuinely 80%-reliable skill still succeeds in three consecutive trials slightly more than half the time. Three wins are compatible with a substantial underlying failure rate. My interpretation is therefore early evidence for a shared execution interface and physical tool adaptation, rather than a controlled demonstration that a fixed simulation-trained policy broadly transfers unchanged. That distinction separates this work’s promise from the transfer objective studied by methods such as DrEureka. Verification becomes a first-class research problem. There is an attractive but dangerous shortcut in the phrase “verifiable code”: treating software validity as physical validity. I would separate three checks. First, does the program satisfy its software interface? Second, do its decisions follow from the state it currently believes? Third, is that state supported by the physical evidence?. A type checker might establish the first. A predicate evaluator might establish the second. Neither, by itself, settles the third. This separation has a useful analogue in task-and-motion planning. PDDLStream, for example, connects symbolic planning with black-box procedures that supply continuous candidates and certify relevant relationships. The interface matters precisely because symbolic structure must be connected to computations about the physical problem. For a coding agent, I would ask a similarly concrete question: what procedure supports each fact that authorizes an action?. A second model reading the same ambiguous image would not automatically satisfy me. I would want to know which errors are shared, which evidence is independent, and what happens when the checks disagree. Abstention is part of that design space. Robots That Ask For Help studies uncertainty alignment for language-model planners, providing a relevant precedent for treating uncertainty as a reason to seek assistance rather than forcing an action choice. In our box example, I would explicitly reward the system for identifying that the current view cannot establish containment. Otherwise, evaluation may unintentionally favor confident guesses over appropriate information gathering. I would also avoid equating final task success with acceptable execution. SafeManip specifically evaluates temporal safety properties in manipulation, moving beyond a terminal-success-only perspective. A useful packaging test should distinguish a careful successful run from one that reaches the same endpoint through prohibited intermediate states. Finally, consider a hypothetical self-training failure. The verifier incorrectly labels a dropped component as successfully placed. The trace enters training. The next model learns from that trace and becomes more likely to repeat the same premature completion pattern. Adding more data would then reinforce the defect. My preferred defense would be to keep the acceptance suite outside the agent’s editing permissions, preserve disagreements as counterexamples, and evaluate verifier changes separately from policy changes. The important property is not that a component is called a verifier, but that its mistakes can be detected without asking the same adaptive loop to certify itself. How I would adapt the idea in a robotics stack. There is an immediate reproducibility boundary. As inspected on September 30, 2026, the public repository contains the report, README, and assets, rather than a runnable release of the full harness and training pipeline. CaP-X provides a public implementation with simulator, robot-interface, and training documentation for related experiments. I would therefore begin with the architecture as a research hypothesis, not assume a turnkey HexaAnything deployment. My first implementation would be deliberately narrow: one existing controller, one task family, and a small set of observable intermediate conditions. I would resist immediately adding a large skill library or autonomous source-code editing. The initial goal would be to determine whether explicit state and verification help at all. For every tool call, I would log the immutable task request, the relevant state snapshot, the tool version, its inputs, the observed result, the verification decision, and the reason for continuing or stopping. Large sensor artifacts could live outside the active model context, provided their references remained recoverable. I would not treat generated explanations as ground-truth labels. If a trace says “the grasp failed because the object slipped,” I would preserve that as a hypothesis unless the retained evidence supports it. Next, I would create a small diagnostic evaluation set. Some trials would contain ordinary execution failures. Others would deliberately challenge state estimation: occlusion after placement, stale observations, swapped object identities, or an action that changes a previously verified relation. The question would be whether the agent notices the right discrepancy and responds appropriately—not merely whether another retry eventually succeeds. For the comparison, I would keep the underlying controller fixed and test three progressively richer conditions: local language instructions alone, instructions with explicit state tracking, and state tracking with verification-triggered recovery. Each would receive a declared observation and computation budget. Only after understanding those differences would I enable persistent memory. I would then ask whether remembered repairs help new instances, rather than simply replaying a successful plan on familiar scenes. Tool rewriting would come later still. Candidate changes would first run against recorded cases and simulation tests, then against a protected regression set. Physical testing would use an approved execution envelope rather than unrestricted agent-generated control. I would keep the candidate and parent versions available side by side. If a change improves deformable-object handling but damages rigid-object performance, the result should be visible as a trade-off, not hidden inside an aggregate score. Finally, I would distinguish the artifact being deployed. Is the deliverable a better controller, a better task program, a better planner checkpoint, or a better verifier? Each deserves its own version, tests, and rollback procedure. This staged approach may sound conservative. That is intentional. It makes it possible to discover whether the representation is doing useful work before investing in an elaborate self-improvement system around it. The durable idea. The most compelling promise of Physical Coding is not that code makes robotics easy. It is that a failure might become a specific, testable change. For me, the decisive question is what remains after the failed episode. Do we retain only a video and a scalar outcome? Or do we retain enough evidence to identify a stale state estimate, an invalid precondition, a poorly specified tool, or a recovery decision worth testing elsewhere?. Those are different learning interfaces. I would not replace capable learned control merely to make a system look more programmable. I would use explicit programs where they make requirements inspectable, uncertainty actionable, and improvements attributable. Then I would demand evidence that the resulting structure pays for itself: fewer false completion claims, better recovery, useful reuse, lower intervention cost, and improvements that survive changes in tasks and physical conditions. That is a narrower promise than unrestricted self-evolving robots. It is also a more actionable research agenda: give the robot a way to check its work, preserve the evidence, and determine whether the next version is actually better.