SpatialClaw: Give the VLM a Geometry Notebook. SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning • Seokju Cho et al. • NVIDIA and KAIST • NeurIPS 2026 / arXiv • 2026. Suppose a mobile robot watches a person walk past a table while its own camera turns. We ask whether the person moved to the table’s left. A vision-language model might answer from appearance. A tool-augmented model might request segmentation, depth, and camera poses. But those measurements do not automatically resolve the question. Someone still has to choose a reference frame, distinguish object motion from camera motion, align observations across time, and decide whether the measurements actually support the conclusion. That “someone” is where SpatialClaw concentrates its effort. The framework gives a vision-language model a persistent Python workspace. Instead of selecting only predefined tool calls—or generating one complete program and hoping it works—the agent writes a cell, executes it, observes the results, and decides what code to write next. Its working materials are images, masks, reconstructed geometry, numerical arrays, and visualizations. The authors present this as a training-free approach to spatial reasoning, rather than a new perception model or a fine-tuned robot policy. The interesting claim is not simply that Python makes arithmetic more reliable. It is that the interface between a model and its tools changes which investigations the model can conduct. Before going further, a provenance clarification matters. The paper was first submitted on June 11, 2026, with Seokju Cho as first author and Min-Hung Chen among its coauthors. I could not retrieve the supplied X post directly. As checked on September 30, 2026, the public paper and NVIDIA’s publication page report a headline improvement of 11.2 percentage points, rather than the supplied summary’s 13.6. This episode follows the released paper, implementation, and tables; the discrepancy remains unresolved. The bottleneck is not necessarily another missing tool. Consider a spatial question that asks which object is closest to a cabinet. A predefined distance tool sounds sufficient—until “closest” means the shortest gap between surfaces, one object is only partially visible, and the best observations come from different frames. Now the agent needs to decide which masks are usable, reconstruct a common coordinate system, select object-associated points, reject obvious contamination, and compute the appropriate distance. It may also need to discover that its first interpretation of the question was wrong. SpatialClaw’s implementation contrasts two ways of restricting that process. Its single-pass configuration requires the model to complete the analysis in one code execution. Its structured-tool configuration instead permits an iterative conversation, but restricts each action to one named tool invocation. The latter retains results in a persistent kernel, so this is not merely a comparison between having memory and having none. Those restrictions affect different dimensions of capability. A single program can contain sophisticated numerical logic. It can branch on whether a mask is empty or iterate over frames. But unless the programmer anticipates the relevant failure, it cannot ask the main model to reinterpret an unexpected visualization halfway through execution. A structured tool agent can inspect results and change its plan. But if its available operations stop at segmentation, reconstruction, and a handful of geometry helpers, an unfamiliar measurement may fall between the predefined functions. SpatialClaw combines iterative observation with open-ended numerical composition. The released environment includes scientific libraries alongside perception tools, allowing later cells to reuse and transform earlier outputs rather than requesting a new high-level API for every spatial relation. The qualification is important: this is not a fundamental limitation of JSON. A JSON tool can launch an interpreter, accept a batch of operations, or expose a rich expression language. Conversely, a Python interface can be so tightly restricted that it offers little additional freedom. The meaningful distinction is the computation the interface permits, the amount of work an action can express, and the evidence returned before the next decision. SpatialClaw’s structured baseline deliberately allows one tool call per step and restricts arguments rather than allowing arbitrary Python expressions. Its results should be interpreted against that concrete alternative, not against every possible structured agent architecture. For roboticists, this reframes a familiar design decision. We often ask which representations and estimators to expose to a high-level planner. SpatialClaw asks another question: after receiving those representations, how much analytical freedom should that planner have?. Inside the reasoning loop. The runtime is organized around a per-example workspace. Input images and metadata are available to the generated code, and intermediate objects survive between steps. The repository also supports separate image sequences for multi-video inputs, rather than silently treating several videos as one continuous stream. A separate planning session runs before execution. It receives the question, metadata, and tool documentation, but not the images. Its role is to outline an investigation, not to answer from linguistic expectations. That distinction is useful: a question can tell the planner that camera motion must be estimated, but it cannot establish which way the camera actually moved. The main model then generates a Python cell. Execution feedback includes printed output, errors, summaries of variables, and images explicitly displayed through. The agent can inspect an overlay, revise a segmentation request, or change a computation before submitting an answer through. An abstract-syntax-tree check screens generated code before execution. Notice the separation between computational state and conversational state. A large point array can remain in the kernel while the model receives a compact summary or a plot. The model does not need every coordinate serialized into its language context to operate on the array. It needs enough information to select the next computation and diagnose whether the previous one was sensible. The feedback implementation is the bridge between those two forms of state. The reconstruction interface supplies depth, camera geometry, and point maps that subsequent code can manipulate. These outputs make it possible to move beyond comparisons in image coordinates and construct measurements in a shared spatial representation. Metric scale still needs an estimator. The underlying Depth Anything 3 nested architecture combines an any-view branch with a metric-depth branch and aligns their scales. That is a learned estimate of scale, not an external measurement certifying every reconstructed distance. Keeping that distinction visible is essential when interpreting answers expressed in meters. SpatialClaw also provides isolated VLM sessions for grounding and visual questions. In the implementation, those sessions use the same configured language-model client as the surrounding workflow. They are additional roles and calls, not evidence of an undisclosed stronger model supplying the geometric answers. The prompt supplies considerable domain discipline. It distinguishes reference frames, encourages robust numerical summaries and magnitude checks, and asks the agent to investigate disagreements between visual and geometric evidence. The main workflow also requests cross-validation before answering. “Training-free” therefore should not be read as “free of spatial engineering.” The engineering has moved into the workspace, instructions, tool contracts, and execution loop. One especially relevant contract concerns frame identity. The implementation uses frame-indexed containers for perception outputs, with alignment checks when compatible objects are composed. A mask and a point map can have perfectly matching array shapes while referring to different moments. Catching that mismatch requires provenance, not ordinary tensor-shape validation. That is a small software detail with a large conceptual consequence. The useful architecture is not unrestricted code replacing structure. It is flexible code operating over evidence whose meaning is partially enforced. What executable spatial reasoning actually buys. A released SpatialClaw example asks for the nearest-point distance between a sofa and a toilet. The reference answer is 2.8 meters; the agent returns 3.143 meters and receives a score of 0.80. The no-tool response returns 4 meters and receives 0.20. This is a useful example precisely because it is not a story of perfect reconstruction: the tool-using answer is better under the evaluation metric, but it is not exact. The project’s trajectory gallery highlights reconstruction, segmentation, scientific-library operations, and repeated visual inspection in this example. These are selected demonstrations, however, not an unbiased estimate of how frequently the system performs each verification successfully. Let us unpack the measurement itself. If we were implementing this investigation, we would first identify usable views of both objects. We would reconstruct compatible geometry, associate image masks with the corresponding point maps, and collect the points belonging to each object. Only then would we calculate the requested relationship. The question asks for the closest surfaces, not the distance between object centers. That distinction can completely change the answer. A long sofa may have a center far from another object while one armrest is nearby. This is where a programmable workspace becomes useful. The analytical procedure can be assembled around the question instead of forcing the question into the nearest available helper function. But numerical execution introduces its own traps. Imagine one erroneous point from the sofa mask landing on the toilet. A minimum-distance calculation could then report almost zero. Now imagine the opposite problem: the closest surfaces are occluded, so the point sets contain only more distant visible surfaces. Even a perfectly computed minimum over those observations would miss the true minimum. Neither error is a Python bug. The response should not automatically be “take the median instead.” A median distance answers a different question. A better investigation would examine why the minimum is suspicious, inspect the contributing points, compare alternative views, and distinguish evidence filtering from changing the target quantity. That is the strongest interpretation of SpatialClaw’s inspect-and-revise design: not that the agent has a universal robust estimator, but that it has an opportunity to choose and question the estimator. The same issue becomes more pronounced for video. Suppose an object shifts left in the image while the camera rotates right. Before classifying the object’s motion, we must decide whether “left” refers to the initial camera, the current camera, the object’s own orientation, or a fixed scene frame. An agent can execute an impeccable subtraction in the wrong coordinate system. Conversely, it can understand the requested frame linguistically but apply the wrong transformation in code. For an embodied system, I would therefore treat the generated program as an inspectable measurement procedure—not as a proof that the resulting spatial claim is correct. Reading the results without flattening the comparisons. The evaluation spans still-image, multi-view, spatial-video, and general-video tasks. The released configuration documentation lists twenty benchmark integrations, including MindCube, MMSI, DSI-Bench, and Video-MME. This breadth is valuable, but it also means the headline combines different kinds of questions rather than measuring a single geometric capability. The protocol caps larger benchmarks at a fixed-seed sample of 1,000 examples. Scoring includes categorical accuracy and numerical mean relative accuracy, with benchmark-specific settings. Consequently, the overall number is an average of benchmark scores—not simply the fraction of all questions answered exactly. The controlled action-interface comparison is the most informative starting point:. | Interface, using Gemma 4-31B | Mean benchmark score | |---|---:| | No-tool baseline | 53.4 | | Single-pass code | 55.2 | | Structured tool calls | 56.7 | | SpatialClaw | 59.9 |. The reported comparison uses the same toolset for the tool-using variants. Read aloud, the story is straightforward. One-shot programming improves on direct answering. Iterative structured tool use improves further. Iterative code achieves the strongest average. But the relevant margins differ. SpatialClaw gains 6.5 points over no tools, 4.7 over single-pass code, and 3.2 over the controlled structured interface. The 11.2-point headline instead compares it with the separate SpaceTools-Toolshed implementation, which scores 48.7 under the shared-backbone evaluation. Those are different experimental questions. The backbone results also deserve careful wording. All six tested backbones improve on average, with gains ranging from 3.1 to 7.7 points. The prominently advertised 59.9 is the Gemma 4-31B result; the table’s Qwen 3.6-27B configuration reaches 62.7. These figures do not establish a simple relationship between parameter count and agent quality. For Gemma 4-31B, prominent gains include 17.6 points on DSI-Bench, 15.3 on MindCube, and 13.4 on MMSI. That pattern is consistent with the hypothesis that compositional spatial analysis helps especially when the answer requires integrating viewpoints or reasoning across time. It is evidence supporting the hypothesis, not a direct measurement of its causal mechanism. Nor is improvement universal: Gemma’s BLINK score falls from 75.7 to 73.4. My practical reading is that SpatialClaw offers a broadly useful inference-time strategy, not a guarantee that more analysis always helps. Every additional perception call, interpretation, and transformation is also another opportunity to introduce error. The deployment question is therefore not just “Does the agent outperform the backbone?” It is “On which questions is the additional computation worth its cost, and when should the system stop?”. What the ablations establish—and what they leave open. The tool ablations are particularly useful because they distinguish the perception models from the convenience functions surrounding them. On the ablation setup, removing utility wrappers changes the average from 56.9 to 56.4. Removing the perception tools produces 51.4, compared with 48.7 for the no-tool baseline. These experiments use a different setup from the headline table: Gemma 4-26B-A4B, fifteen benchmarks, and up to 500 examples per benchmark. The near-maintenance of performance without utility wrappers supports an important engineering proposition: a capable model can reconstruct some missing helper logic using general numerical libraries. You may not need to anticipate every geometric operation in your API. I would not conclude that wrappers are unnecessary. In a deployed system, a tested helper can enforce units, numerical conventions, runtime limits, and error behavior. Small differences in benchmark score do not capture those benefits. Similarly, the no-perception result suggests that the remaining agent machinery contributes something beyond specialized segmentation and reconstruction. But it is not a pure test of Python syntax. The framework still offers an iterative workflow, numerical computation, and visual access. Several mechanisms change together. The single-pass comparison has another qualification: its configuration disables planning and collapses execution to one step. Thus, the difference from SpatialClaw does not isolate intermediate visual feedback alone. It compares two complete operating modes. The structured baseline is more informative about composition because it retains iterative interaction and persistent results. Nevertheless, equal step counts do not imply equal computation. One code cell can combine several tool calls and numerical operations, while one structured action invokes a single named tool. For a follow-up study, I would want performance curves against several budgets: total model tokens, perception calls, GPU time, and wall-clock latency. I would also separate persistent state, visual feedback, planning, and action batching. That would help answer whether the advantage comes primarily from better investigations, more work per turn, more opportunities to recover, or some combination. There is another distinction worth testing directly: verification instructions versus verification effectiveness. A prompt may request two supporting lines of evidence, yet both can inherit the same erroneous reconstruction. A point-cloud visualization and a distance printed from that point cloud are not independent measurements of scale. A stronger experiment would deliberately introduce inconsistent masks, corrupted calibration, or ambiguous reference frames and measure whether the agent detects the problem, repairs it, abstains, or merely produces a more elaborate justification. For embodied applications, that recovery profile may matter more than another small increase in average score. Where SpatialClaw sits among related approaches. SpatialClaw belongs to an established line of work on executable agent actions. Xingyao Wang and colleagues’ CodeAct already proposed Python as a unified action space, with multi-turn execution and revision based on observations. That precedent matters: SpatialClaw should not be understood as inventing the general idea of a language model operating through an interpreter. Its distinctive emphasis is spatial evidence: how a model combines perception outputs, numerical geometry, visual inspection, and question-specific computation. pySpatial is a close domain-specific neighbor. Zhanpeng Luo and colleagues use generated Python programs to compose reconstruction, camera-pose recovery, and novel-view rendering. Their framework is also zero-shot, and they report real-world indoor navigation experiments in addition to benchmark evaluation. Python-based, training-free spatial reasoning therefore predates SpatialClaw as well. The relevant comparison is how the agent can revise its investigation, not whether one approach has geometry and another does not. SpaceTools addresses a different dimension: learning tool coordination. Siyi Chen and colleagues introduce Double Interactive Reinforcement Learning, combining a teaching phase with subsequent interactive exploration. Their work includes real-world manipulation using a seven-degree-of-freedom robot as a tool. This context prevents an easy misreading of SpatialClaw’s headline. Evaluating another framework with the same substituted Gemma backbone is useful for studying portability and inference-time behavior. It is not equivalent to comparing every method’s original trained system under its intended conditions. My view is that these approaches are more complementary than their leaderboard rows suggest. CodeAct supplies a general action-space precedent. pySpatial explores explicit 3D visual programming. SpaceTools studies how to train coordination. SpatialClaw makes a focused case for iterative, stateful code composition across a broad spatial evaluation. A natural next step would combine a flexible analytical workspace with training that rewards correct tool selection, geometrically valid programs, and successful recovery from misleading evidence. How I would adapt it for an embodied stack. I would begin by putting SpatialClaw-style reasoning on the read-only analysis side of a robot system. Let it inspect observations, construct measurements, identify inconsistencies, and propose spatial constraints. Do not begin by giving generated Python direct authority over actuators. That is a deployment recommendation, not a robot-control result demonstrated by SpatialClaw. The repository’s architecture separates three services: VLM inference, GPU-backed perception, and the agent process managing Jupyter kernels. SLURM orchestration is provided, but the services also have direct entry points. This separation offers a practical starting point for integrating the reasoning component without merging it with an existing control process. For reproduction, I would first use the released experiment paths rather than immediately replacing half the stack. The running guide exposes , , and executor modes for the interface comparison. Establishing that baseline gives subsequent changes—new sensors, different reconstruction, smaller backbones—a reference point. Then I would strengthen the evidence contract. A robot-facing spatial answer should include its reference frame, units, observation time, relevant object identities, and the evidence used to derive it. If it says an obstacle is forty centimeters away, the consumer needs to know whether that means nearest observed surface, object centroid, or a projected ground-plane distance. I would also require uncertainty or an explicit statement that uncertainty is unavailable. A decimal-valued answer should not gain authority merely because the program printed three digits. SpatialClaw’s frame-indexed containers provide a useful precedent for this approach. I would extend that idea to coordinate-frame identity, calibration version, and units, so that more semantic mistakes become detectable contract violations rather than plausible numbers. Next comes execution isolation. The released framework includes static code screening, but I would treat that as one layer rather than permission to expose a robot’s host environment. My adaptation would put kernels behind process or container boundaries, restrict filesystem and network access, limit resources, and keep actuator interfaces outside the generated-code environment. The static checker is a useful precaution, not the whole deployment design. I would preserve replayable artifacts as well. The important record is not only what the model said it intended to do. It is which frames entered each tool, which arrays were produced, which transformations were applied, and which displayed evidence prompted a revision. This creates a useful division of labor. The generative component can propose an unfamiliar computation; deterministic checks can verify its frame usage, units, numerical validity, and compliance with a task contract. The evaluation would then move beyond generic question answering. I would build a task-specific set containing occlusion, reflective surfaces, small clearances, moving cameras, repeated-looking objects, and stale observations. I would measure failure detection and abstention alongside answer quality. I would also compare against a simpler fixed pipeline. An agent is valuable when the investigation genuinely varies. If nearly every query needs the same calibrated operation, a tested deterministic implementation may be faster, easier to validate, and more predictable. Finally, account for infrastructure and licensing before treating this as a lightweight plug-in. The installation guide documents GPU and model-access requirements, and the repository is released under the NVIDIA Source Code License-NC. “Released code” should not be casually translated into unrestricted commercial availability. The engineering objective is not to maximize the amount of generated code. It is to give the model enough freedom to investigate unusual situations while making its conclusions easier for the rest of the robot stack to challenge. The takeaway: make spatial claims inspectable. SpatialClaw invites us to reconsider where spatial competence should reside. Some of it belongs in learned perception. Some belongs in numerical geometry. Some belongs in the model’s interpretation of the question. And some belongs in the software interface that determines whether those components can correct one another. My strongest takeaway is that a spatial reasoning agent should be able to change its measurement procedure after seeing the evidence—not merely change the wording of its answer. For research, that suggests treating action expressivity, feedback design, and semantic data contracts as first-class variables. For robotics, it suggests an analytical component that can construct useful spatial claims while leaving their validation and physical consequences under explicit control. The goal is not a model that sounds more certain about three-dimensional space. It is a system whose spatial conclusions can be inspected, reproduced, challenged, and, when necessary, rejected. Source scope: the released materials support the baseline-specific results discussed here; they did not provide enough information to reconcile the supplied announcement’s +13.6 figure.