We present a framework for assistive robot manipulation that addresses two fundamental challenges: efficient adaptation of large-scale models for scene affordance understanding and effective learning of robot actions by grounding...
Stay in the loop on research in AI and physical intelligence.
Teach the Robot Where to Act—Then Make It Fast.
Affordance-based robot manipulation with flow matching • Fan Zhang and Michael Gienger • Honda Research Institute EU, Offenbach, Germany • Frontiers in Robotics and AI • 2026.
Two efficiency problems, one manipulation pipeline.
Imagine a robot looking at a table containing a cup, a brush, a utensil, and some tissues. A person is sitting nearby. The instruction changes, but the image does not.
One request makes the brush handle and the person’s head relevant. Another makes the utensil, food, and mouth relevant. Recognizing every object is not sufficient. The robot needs to identify which locations matter for this particular interaction—and then generate a movement that connects them appropriately.
Zhang and Gienger’s proposal is to make that intermediate reasoning explicit. A language-conditioned perception module predicts manipulation affordances, and a flow-matching policy converts the resulting visual conditioning into robot actions. The project therefore connects two questions that are often investigated separately: how to adapt pretrained perception economically, and how to generate actions with low inference latency.
Those are different kinds of efficiency. Updating fewer parameters does not automatically make perception cheap to execute. Generating an action sequence quickly does not automatically make a robot responsive. And producing a useful intermediate representation does not establish that the complete system generalizes beyond its collection environment.
That separation is the key to reading this work productively.
The journal article was published on September 22, 2026; the original preprint appeared on September 2, 2024. It is best understood as a focused investigation of an architectural combination, rather than a claim that flow matching has only just arrived in robotics.
For practitioners, the interesting question is not simply whether to replace diffusion with flow matching. It is whether explicit spatial supervision can make the downstream action problem easier enough that a small number of generation steps becomes useful—and whether that benefit survives realistic deployment conditions.
Affordances as an interface, not a complete plan.
The paper represents affordances with image-space heatmaps. In a feeding example, relevant regions include a utensil handle, the food, and the mouth. The action model receives visual information alongside those maps; for three-dimensional trajectory prediction, the conditioning includes RGB-D. This is therefore an added spatial guide, not a replacement of the scene with a few isolated coordinates.
Consider why that distinction matters. A heatmap can indicate where to grasp without representing the full geometry of the grasp. It can identify a destination without saying whether the approach should come from above or from the side. It can mark two interaction regions without resolving the order in which contacts should occur.
My interpretation is that this interface reduces the policy’s search burden rather than completing its reasoning. The downstream model still has to learn how image geometry, robot motion, and task progress fit together.
There is a useful comparison with Wentao Yuan and colleagues’ RoboPoint. RoboPoint predicts image keypoint affordances from language and images, using an automatic synthetic-data pipeline to develop spatial grounding. Its central output is an actionable location representation, rather than a complete manipulation controller. That distinction helps separate the quality of a spatial interface from the quality of the policy consuming it.
For example, suppose a perception system highlights the correct cup handle but the robot approaches from an infeasible direction. The spatial prediction may be right while the manipulation fails. Conversely, a policy might complete a task despite a slightly displaced heatmap because the surrounding image contains enough corrective information.
I would therefore avoid evaluating an affordance interface only by asking whether its brightest pixel matches an annotation. The more consequential question is how downstream success changes as the map becomes incomplete, ambiguous, or slightly wrong.
The study’s ADL dataset contains 10,000 demonstrations across ten tasks, collected through kinesthetic teaching, with approximately thirty objects and a manikin-centered setup. The reported train/test split is random, at eighty/twenty.
That supports a controlled manipulation study. My proposed generalization test would be stricter: hold out collection sessions, object instances, viewpoints, or entire scene layouts. A random split answers whether a model predicts unseen examples from the collection process. A grouped split asks whether it survives a meaningful change in that process. For an assistive system, I would want both answers before treating a good test score as evidence of portability.
Prompt tuning changes processing without changing backbone weights.
The perception strategy belongs to the family established by Menglin Jia and colleagues’ Visual Prompt Tuning. Instead of updating a pretrained visual backbone, VPT introduces trainable tokens into its input stream. The shallow variant introduces prompts near the beginning; the deep variant introduces prompts at multiple transformer layers. The backbone weights remain fixed while the added prompts steer its computation.
The important mechanism is not that a frozen network somehow becomes trainable again. Its parameters remain unchanged. But attention depends on the tokens presented to it. Changing those tokens can change the features produced by the same fixed transformation. Gradients still have to propagate through relevant computations to optimize the prompts.
Zhang and Gienger use a frozen ViT-B/16, text-conditioned prompts, and a lightweight transformer decoder. Their deep configuration injects language-derived features at successive vision layers; the shallow configuration introduces them only at the first layer. Training updates the prompt-related components and decoder.
I would describe this as learning how to interrogate a visual representation. The adaptation problem becomes: which language-conditioned signals cause the existing visual computation to expose the information needed for this manipulation task?.
That is a more precise description than “the model understands the instruction.” It makes the mechanism inspectable without assuming that task-relevant activations imply broad semantic competence.
The parameter counts also deserve attention. Deep prompt tuning uses 42.1 million trainable parameters versus 153.8 million for full tuning. Heatmap-center errors are 2.93 versus 1.15 pixels; AffordanceLLM reaches 1.01 pixels.
My reading is a trade-off, not an across-the-board accuracy win. The reduction in trainable parameters is substantial, but this is not adaptation using only a negligible handful of vectors. Nor do fewer trainable weights establish an equivalent reduction in deployment memory, wall-clock training cost, or inference latency.
If I were adapting this idea, I would separately measure optimizer-state memory, activation memory, training throughput, and end-to-end perception latency. I would also account for annotation effort. Saving GPU memory while requiring additional interaction labels may be an excellent exchange—but it is still an exchange.
The experiment I would most like to add is an annotation-budget comparison: how much downstream manipulation performance can each adaptation method obtain per hour of demonstration collection and labeling?.
Flow matching: learn how action samples should move.
The action-generation component starts from the flow-matching formulation introduced by Yaron Lipman and colleagues. Instead of learning a direct observation-to-action regression, the model learns a time-dependent vector field that transports samples from a simple distribution toward a data distribution. Conditional flow-matching objectives make that training practical without simulating the complete transport process inside every optimization step.
For robot behavior, the sample being transported can be an entire action sequence. Think of a rectangular array: future timesteps along one axis, action coordinates along the other. Generative inference transforms that array into a plausible action chunk.
The released implementation uses TorchCFM’s independent conditional-flow construction. In plain language, take a demonstrated action array and an independently sampled noise array of the same shape. Select a point along the straight interpolation between them. The supervised target is the displacement from the noise array to the demonstrated array. The network learns to predict that displacement direction from the interpolated sample, generation time, and observation conditioning.
This is ordinary supervised regression on a carefully constructed target. The difficult distribution-learning problem is expressed through many simple training examples.
The distinction between the training target and the final generator is worth retaining. During training, the construction knows both endpoints. At deployment, the network does not know the demonstrated endpoint corresponding to its initial noise. It must produce a useful vector field from the observation and the evolving sample.
The authors’ Push-T example makes the implementation concrete: Gaussian training noise, a conditional flow matcher with zero additional path noise, and mean-squared error between predicted and target vectors. Observation features combine a ResNet-18 image representation with the end-effector position.
The temporal model is a conditional one-dimensional U-Net. Its convolutions operate over the action sequence, while feature-wise linear modulation supplies global conditioning. The implementation combines the observation condition with the generation-time embedding inside the action network. This is recognizably close to the architecture many practitioners already use for Diffusion Policy.
That architectural continuity is useful experimentally. If you already have an action-chunking policy, you can investigate the flow objective without simultaneously changing every other component. My preferred first experiment would preserve the encoder, action representation, normalization, data loader, and evaluation environments.
Otherwise, an apparent gain from flow matching could actually come from a better encoder, a different horizon, or a more favorable normalization scheme.
Normalization is especially easy to underestimate. The released Push-T data utilities scale coordinates using dataset minima and maxima into a common range.
For a new robot, I would explicitly inspect the relative numerical scales of positions, rotations, joint coordinates, and gripper commands. The regression objective sees those scales. Before tuning a generative model, I would want to know whether the loss is emphasizing the dimensions that matter physically.
Generation time is not robot time.
There are two clocks in an action-chunk generator.
One clock indexes the future actions the robot might execute. The other indexes the numerical procedure that creates those actions. A generation update changes a candidate action array; it is not itself a physical command to the robot. Flow matching concerns the latter process. Diffusion Policy likewise separates iterative generation from receding-horizon execution.
This distinction prevents several misleading interpretations.
A straight interpolation in the generator does not mean the end effector should move along a straight Cartesian line. A single generation step does not mean the robot completes the task in one control cycle. And a smoothly varying generative vector field is not, by itself, a statement about acceleration or torque along the executed motion.
It is also important not to turn the comparison into “deterministic flow matching versus stochastic diffusion.” Ruiqi Gao and colleagues’ Diffusion Meets Flow Matching explains how Gaussian flow matching and diffusion formulations can be related through parameterization and schedule choices. DDIM can generate deterministically once its initial noise is fixed. Their analysis also distinguishes straight conditional interpolation paths from potentially curved sampling trajectories.
So my takeaway is narrower than a categorical victory for one model family. The useful empirical question is whether a particular objective, representation, and sampler achieve the required action quality with fewer network evaluations.
Nor is fast generation unique to an undistilled flow model. Prasad and colleagues’ Consistency Policy distills a pretrained diffusion policy into a fast student. Wang and colleagues’ One-Step Diffusion Policy similarly targets single-step action generation through diffusion distillation. These are different training workflows, but relevant alternatives when deployment latency is the real constraint.
I would also distinguish successful one-step control from faithful multimodal sampling. If observations make most decisions nearly unambiguous, a low-diversity predictor may perform well. That does not establish that it preserves multiple valid strategies when the task genuinely branches.
AdaFlow offers a useful comparison here: it adjusts integration effort using a learned variance-related signal, spending more computation where generation is difficult and approaching one-step behavior in low-variance cases. That suggests testing where few-step inference works, rather than assuming one fixed budget is equally appropriate everywhere.
What the numerical results actually establish.
The trajectory-prediction result is the clearest efficiency finding. On an RTX 4090, Table 4 reports one-step flow matching at 1.031 centimeters error and 13.737 milliseconds, versus sixteen-step DDIM at 2.265 centimeters and 105.583 milliseconds. The prose gives slightly different timing values.
Using the table values, I calculate approximately an eighty-seven percent latency reduction. I would report that as a substantial result for this tested configuration, while avoiding false precision when summarizing the article’s slightly inconsistent timing descriptions.
More importantly, I would keep the comparison attached to its operating point: hardware, model, task distribution, action representation, and sampler budget.
Suppose your application can tolerate a hundred milliseconds of action-generation latency but cannot tolerate a particular grasp failure. The fastest generator is not necessarily the best choice. Conversely, suppose target motion makes that latency unacceptable. A modest quality difference may become secondary to whether the robot receives a sufficiently current command.
Those are deployment decisions that a single offline error metric cannot resolve.
The real-robot evaluation provides a second piece of evidence: fifty trials per method produce success rates of 82% for affordance-conditioned flow matching, 76% for diffusion, 44% for transformer behavior cloning, and 74% for end-to-end flow matching. The principal flow result uses sixteen inference steps.
Two cautions follow from those numbers.
First, the hardware result should not be silently attached to the one-step timing headline. They are different experimental conditions. I would want a hardware sweep over generation budgets before concluding that the fastest configuration retains the same success rate.
Second, the comparison with end-to-end flow is suggestive rather than definitive. Across fifty trials, the reported eight-percentage-point difference corresponds to four additional successes. That is worth investigating, but I would not treat it as a precise estimate of a broadly generalizable advantage.
There is also a supervision question. To isolate the value of the intermediate representation, I would compare models with matched access to annotations. An auxiliary affordance-prediction objective attached to an otherwise direct policy would be one useful control. It would help distinguish the benefit of extra spatial supervision from the benefit of explicitly routing that supervision through heatmaps at inference.
The simulation evidence should be read separately. The benchmark implementation includes Push-T’s image-conditioned controller, while the Kitchen example uses low-dimensional state inputs. These are tests of action-policy learning, not equivalent demonstrations of the complete language-to-affordance pipeline.
The repository reports maximum and average evaluation scores, which should not be conflated. Push-T also uses target coverage rather than a directly interchangeable binary manipulation-success metric. Its environment computes overlap between the block and goal geometry.
A useful counterexample to a universal-performance narrative appears in Kitchen: DDIM’s reported average is 0.7471, slightly above flow matching’s 0.7425.
My synthesis is that the action-quality results are competitive, while the reduced-step operating points are the more interesting practical contribution. I would not select a policy family from the largest number in a mixed benchmark table. I would select it from a quality-versus-latency curve measured on the target system.
Fast inference is not yet closed-loop assistance.
The ADL setup predicts a complete thirty-two-waypoint trajectory and executes it open-loop, with gripper closure prescribed at the first waypoint. In contrast, the public benchmark policies use receding-horizon execution.
This difference is central.
Imagine that a tool slips after grasping. A fast trajectory generator helps only if the system observes the slip, represents the new situation, invokes the policy again, and can replace the pending actions appropriately. A low inference time is one component of that response. It is not the response itself.
The Push-T example illustrates the distinction concretely: it predicts sixteen actions and executes eight before replanning.
I would therefore profile more than neural-network execution. My deployment trace would timestamp image acquisition, preprocessing, affordance prediction, action generation, command transmission, controller acceptance, and the next observation. I would also record how many already-buffered actions can still execute after a new plan becomes available.
That trace would answer a more useful question than frames per second: how old is the information governing the next physical movement?.
The article acknowledges that its model does not explicitly enforce second-order smoothness or torque feasibility.
For my own implementation, I would preserve an independent execution layer responsible for motion limits, feasibility checks, and stopping behavior. I would not infer those properties from the generator’s mathematical smoothness or from successful nominal demonstrations.
That is especially important when interpreting the assistive motivation. A controller that moves a tool toward an annotated region is not automatically a validated system for interacting safely with a moving person. My next experiments would begin with object and manikin perturbations, establish recovery behavior there, and only then consider more demanding interaction studies.
Where this sits among neighboring approaches.
PointFlowMatch, by Eugenio Chisari and colleagues, investigates conditional flow matching with point-cloud observations and studies action-representation choices, including rotation geometry. It provides a useful contrast: rather than emphasizing a language-conditioned image-space interface, it emphasizes the geometric information supplied to the policy.
I would use that comparison to decide where the main uncertainty lies in a new project. If the system consistently selects the wrong interaction region, explicit semantic grounding may be the priority. If it selects the right region but fails under viewpoint changes or unfamiliar orientations, I would investigate geometric representations before assuming that a different action generator will solve the problem.
At a different scale, Physical Intelligence’s π₀ combines a pretrained vision-language model with flow-based action generation for general robot control. It shows that flow matching can also be used without this paper’s particular explicit affordance interface.
These comparisons should not be converted into an imaginary leaderboard. The methods differ in data, embodiment, representation, and training scale.
The design question I take from them is where to place structure. You can put more structure into the perception interface, more geometry into the observation representation, or more broad prior knowledge into a large pretrained policy. Zhang and Gienger’s work is useful because it makes one of those choices unusually visible and testable.
Reproducing the idea: start with the artifact boundary.
The authors provide the repository. Its README advertises Push-T and Franka Kitchen training/evaluation examples, links pretrained benchmark weights, and acknowledges dependencies on Diffusion Policy and TorchCFM. That is a useful starting point for investigating the action model. It should not be mistaken, without further checking, for a turnkey reproduction package for every ADL experiment.
My first step would be to reproduce one public benchmark before adding a new affordance stack. That keeps the debugging problem small. If action normalization, checkpoint loading, or horizon alignment is wrong, adding language-conditioned heatmaps will only make the failure harder to localize.
There is a concrete implementation detail worth auditing: the released Push-T example uses for training noise but for test initialization—a Gaussian/uniform mismatch. This is a code-inspection finding, not evidence about which implementation produced the published tables. I would reconcile it before calling a run a faithful replication.
I would also lock the software environment. The requirements file leaves several important dependencies unpinned, including PyTorch, Diffusers, and TorchCFM.
For research code, I would record the repository revision, package versions, dataset preprocessing, normalization statistics, and checkpoint-selection rule together. A pretrained weight file without those surrounding decisions is not a complete experimental specification.
Next, I would establish two separate timing measurements. One would time the action generator with conditioning already available. The other would time the complete observation-to-command path. Their difference would tell me whether optimizing generation is actually the highest-value engineering task.
Finally, I would add the affordance model only after defining its contract. What coordinate frame does it use? Can it produce multiple plausible interaction regions? How are missing detections represented? Does the policy receive predicted maps during training, or only clean annotations? Those are the questions I would resolve before tuning the combined system.
This staged approach is less glamorous than replacing the entire stack at once. It is also more likely to reveal which component is responsible for an improvement.
The experiments I would run next.
The most interesting unresolved question is whether explicit affordances specifically make few-step action generation easier, rather than merely improving perception.
I would test that with a crossed experiment: raw observations versus affordance-augmented observations, each paired with diffusion and flow objectives. I would hold the encoder capacity, action representation, training budget, and data access as constant as possible. Then I would sweep generation steps and compare the resulting latency-quality curves.
If the affordance-conditioned flow model retained performance at a lower generation budget than the corresponding raw-observation model, that would support the proposed interaction between representation and sampling efficiency.
My second experiment would deliberately perturb the interface. I would displace heatmap centers, broaden them, remove one task-relevant region, and introduce a competing region. I would measure not just whether performance drops, but how it drops.
An abrupt failure under tiny map errors would suggest brittle reliance on the intermediate prediction. Gradual degradation, or recovery using the remaining visual evidence, would be more encouraging. I would not assume that an interpretable interface is automatically a robust one.
Third, I would test recovery rather than only nominal execution. Move the object after the initial observation. Interrupt a grasp. Change the target pose after an action chunk has been generated. Ask whether the system reobserves, revises its plan, and completes the task—not merely whether its next sampled trajectory looks plausible.
Fourth, I would separate linguistic variation from task novelty. Paraphrasing a known request is a different test from combining familiar objects in a new interaction. I would evaluate those conditions independently, rather than putting them under one broad generalization label.
Finally, I would inspect multimodality directly. For an observation with several valid strategies, I would sample repeatedly and ask whether the outputs remain valid, meaningfully diverse, and temporally consistent. A low average distance to one demonstration would not be my only criterion.
These experiments would move the discussion from “the architecture works on this collection” toward a clearer account of when its intermediate representation helps, when few-step generation is justified, and where the complete system remains vulnerable.
The practical takeaway.
What I would take from this paper is not a universal instruction to replace diffusion, and not a claim that heatmaps solve assistive manipulation.
I would take a modular hypothesis worth testing: make task-relevant interaction locations explicit, then learn an action generator that can exploit that structure efficiently.
The strongest engineering habit suggested by the work is to keep the claims separate. Measure adaptation cost separately from inference cost. Measure trajectory prediction separately from task completion. Measure action-generation latency separately from feedback responsiveness. And distinguish successful nominal behavior from recovery under a changed scene.
For a team already using action-chunking imitation learning, that is a tractable research program. Start with a controlled flow-matching baseline. Add spatial supervision where it addresses an identifiable failure. Then test whether the combination actually improves the quality-latency trade-off on the robot.
The promise is a more inspectable route from task intent to action. The next challenge is showing that the route remains useful when the world changes after the robot starts moving.