🤗 Upvotes: 254 | cs.AI, cs.CL, cs.GT, cs.MA, cs.NE
Authors:
EverMind AI
Title:
Raven: The Harness of Harnesses for Composable Agentic Intelligence
Arxiv:
http://arxiv.org/abs/2609.33439v1
Abstract:
As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to specific domains limits the generality of a single harness. The central question thus shifts from how to engineer a stronger harness for one domain to how to autonomously construct specialized harnesses, improve them through experience, and orchestrate them across domains. We introduce Raven, \emph{The Harness of Harnesses}, an open-source multi-agent ecosystem that automatically constructs and evolves modular harnesses for specific models and domains, treating each executable model--harness pair as a composable unit of intelligence. To support an \emph{All-Domain Collaboration Network}, its Host Agent decomposes goals, matches subtasks to specialized agents, coordinates execution dependencies, and integrates results, while a host archive and EverOS preserve experience across tasks and Skill Forge makes that experience available as reusable procedures. Our theory establishes sufficient conditions for such composition to expand reliable task coverage beyond that of the available individual agents under a shared resource budget. On complex and long-horizon tasks, Raven significantly outperforms the state-of-the-art agent systems, pushing the frontier of composable agentic intelligence.
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.com
Creator:
Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/
Gengyu Wang, LLM ML, http://wanggengyu.com
Listen on:
Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXL
Apple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236
Cover Image by Kawen Kuang https://kawen.art
Evan: Welcome to Daily Paper Cast.
Evan: Today's paper is from the Hugging Face daily paper list of September 30, 2026, with 254 upvotes.
Ashley: The title is Raven: The Harness of Harnesses for Composable Agentic Intelligence.
Ashley: It's authored by Chuanrui Hu and Dizhan Xue, with Chuanrui Hu as the corresponding author, representing EverMind AI.
Evan: Let's dive into the Introduction.
Advances in large language models, including instruction following and code generation, have provided a foundation for agents that interpret user goals and act through software.
Ashley: Indeed, language models can even learn to invoke external tools.
Agent systems organize these capabilities into multi-step interactions with external environments.
This includes reasoning combined with actions and observations, which allows agents to gather information and adapt their decisions based on feedback.
Evan: Specialized execution interfaces shape what agents can accomplish.
For example, in automated software engineering, these developments have shown how agents with different expertise and execution mechanisms can work toward a shared goal.
Ashley: Complex goals often require several specializations within one workflow.
For instance, developing and operating a first-person shooter game involves requirements research, gameplay programming, visual design, integration testing, and sustained operation.
Each of these stages involves different tools and criteria, and the outputs of one stage support the next.
Evan: So, an agent’s capability depends not only on its model but also on its harness—the tools, context management, skills, execution policies, and recovery mechanisms around the model.
These mechanisms dictate how an agent applies its model’s capabilities within a domain.
Ashley: Exactly, and this leads to two main challenges.
First, as we support more tools and execution conditions, the number of interacting design choices increases, making manual harness development challenging to scale across various models and domains.
Evan: And second, specialized harnesses often encode assumptions about their tools, working context, and expected outputs.
Combining capabilities requires careful matching of subtasks to suitable executors and ensuring compatible handoffs.
Multi-agent failures often stem from inter-agent misalignment and inadequate task verification.
Ashley: Mechanisms like multi-agent conversation and role-based workflows help to organize collaboration.
However, the benefit of such compositions depends on task structure, local capabilities, and the cost of coordination.
The central problem then becomes constructing and improving specialized harnesses while ensuring their composability across different domains.
Evan: To address these challenges, the paper introduces Raven, The Harness of Harnesses, an open-source multi-agent ecosystem.
This system treats each executable model–harness pair as a unit of composition.
Ashley: Raven includes the native specialists Raven-Research, Raven-Code, Raven-Design, and Raven-Oncall, as well as independently developed agents like Claude Code, Codex, and Hermes Agent.
Execution adapters connect these agents to a shared orchestration interface while preserving their tools and internal execution policies.
Evan: Within this system, the Host Agent organizes collaboration by matching subtasks to registered capabilities and mapping these dependencies as a directed acyclic graph or DAG.
Each node represents an agent, and the edges denote dependencies between their outputs and subsequent tasks.
Ashley: Moreover, Raven combines harness adaptation with the reuse of experience across tasks.
Whereas prior methods adapt agents by retaining insights from past executions or distilling interactions into reusable skills, Raven modular harnesses expose execution policies that external evolvers can modify within defined boundaries.
Evan: Building on previous work, HarnessBank, the evolution process diagnoses failures, proposes candidate harnesses, and evaluates their behavior with a fixed task model.
Persistent memory complements these policy changes with a host archive retaining user context and an optional EverOS backend providing semantic access to past experiences.
Ashley: Furthermore, Skill Forge allows for procedural reuse by retrieving task-relevant procedures from local skills, memory-derived skills, and SkillHub, a corpus and retrieval design based on prior work called SkillCorpus.
Evan: In addition to the system design, the theoretical analysis in the paper characterizes when composition can expand the capabilities of the available agents.
For a specified agent pool and task family, it formalizes capability as reliable task coverage under a common resource budget.
Ashley: To dive into specifics, the theory establishes sufficient conditions under which complementary local capabilities, compatible handoffs, and bounded planning and execution errors allow the composed system to solve tasks that the individual agents cannot reliably solve alone under the same budget.
Evan: Essentially, Raven aims to improve task performance on complex and long-horizon tasks, significantly outperforming state-of-the-art agent systems.
This pushes the frontier of composable agentic intelligence.
Ashley: Indeed, and this concludes the Introduction section of the paper.
Evan: Alright, let's dive into the methods detailed in the paper.
Ashley: The Method section covers several key components of the Raven ecosystem, including its architecture for multi-agent collaboration, harness self-evolution, memory systems, and skill retrieval processes.
Evan: Starting with multi-agent collaboration.
Raven's design includes a Host Agent that decomposes tasks, matches subtasks to specialized agents, coordinates execution dependencies, and integrates results.
Ashley: Exactly.
The Host Agent uses an orchestration guidance tool which helps to compose a graph for complex tasks.
It starts by loading the orchestration guidance on demand when a request needs multiple agents.
This is done to avoid incurring an input cost for each turn that doesn't require orchestration.
Evan: Once the host determines that a request requires orchestration, it loads the full set of rules necessary to construct a graph.
These rules specify node wiring, placeholder syntax, concurrency limits, and whether single delegation is preferable.
If the guide is missing, the system provides feedback to direct the host to relevant rules.
Ashley: After constructing the graph by working backward from the requested deliverable, each node is assigned a unique identifier and an executor able to produce the output.
The Host Agent then adds dependencies indicating which node's output is required as input for another node.
Evan: Admission validation follows.
The runtime admits the graph after five groups of checks: format, graph structure, agent capability, agent status, and environment.
Ashley: These checks ensure the nodes are well formed, dependencies are resolvable, paths are confined within permitted roots, agents are registered and enabled, and that the process does not violate dispatch budget or recursion rules.
Validation stops at the first error and returns a message naming the offending node or field, often suggesting a repair.
Evan: Once the graph is admitted, each node proceeds through the runtime, moving through states like pending, running, exception, completed, failed, skipped, or canceled.
Nodes run once all dependencies are settled, and output or raised errors return to an independent model call judge.
Ashley: The judge must validate the completion of nodes and either accept them as accomplished or mark them as exceptions for host intervention.
Clarification requests are sometimes made by workers, seeking information omitted in the prompt.
The host can attempt to answer these from its current conversation and memory or forward them to the user if necessary.
Evan: After execution, a run report is generated summarizing intermediate outcomes, including completed nodes, failures, cancellations, skips, node output files, and terminal outputs.
Reports enable the host to synthesize the final deliverable from the nodes' outputs.
Ashley: Nodes are key here.
Artifacts and memory flow separately; artifact completion releases dependencies, while memory recording proceeds asynchronously.
File-based node records allow successors to read upstream outputs inline or by reference paths, without duplicating artifacts in prompts.
Evan: Cross-graph reuse is supported with node identifiers and shared context blocks, allowing later runs to reference earlier outputs without executing the node again.
This reduces redundant subtasks and limits host context references.
Ashley: Group memory is also notable.
Shared group memory converts observations into assessments that inform future orchestration and worker invocations.
Round summaries, agent assessments, and user feedback provide a collective record that informs planning and retrieval during subsequent tasks.
Evan: This memory-management strategy ensures persistence, context transfer, and execution synchronization, supporting effective agent collaboration over long periods and multiple tasks.
Ashley: Now, let's move on to the second component: harness self-evolution.
Raven's method for harness adaptation builds on HarnessBank.
It exposes the model–harness pairs through modular strategy interfaces, evaluates their performance, diagnoses failures, and evolves harnesses by proposing candidates and screening them under fixed protocols.
Evan: Failures are monitored, and new candidate harnesses are proposed to address diagnosed issues.
These candidates undergo a strict screening process evaluating validity, activation, and paired gain, retaining only the strongest performers for further evaluation.
Ashley: Harness candidates are categorized by the types of modifications they include—prompt, knowledge, runtime, and configuration edits.
Selected candidates are then stored in a Harness Gene Bank for reuse and further improvement.
Evan: A Task Agent executes tasks, an Evolver Agent proposes modifications, and a fixed Evaluator measures outcomes.
Modular harness design allows for changes without affecting the underlying model, focusing on improving how the agent applies its model capabilities within domains.
Ashley: Evaluation uses a fixed protocol specifying model versions, task sets, and environment conditions.
Successful interventions are validated for their causal contribution to performance improvements.
Complex tasks guide the selection process but do not provide family-wise guarantees.
Evan: Memory systems help agents preserve interaction histories and procedural knowledge.
Episodic trace formation groups interactions into segments, which undergo semantic consolidation into MemScenes, identifying trends and persistent information.
Ashley: Reconstructive retrieval recalling episodes and facts uses BM25 and dense search combined through reciprocal rank fusion.
Skill Forge retrieves procedures, updates skills from execution cases, sorts reusable insights, and appends validated skills to the memory store.
Evan: The retrieval ensures agents have access to prior experiences to inform subsequent tasks.
Each agent session records useful methods and fixes to maintain continuity and adaptability across tasks.
Ashley: That concludes the Method section.
Evan: Let's move on to the Experiment and Results section to understand how Raven performs.
Ashley: Yes, the paper evaluates Raven's effectiveness in organizing specialist work, improving execution with a fixed model, and reusing procedural knowledge.
Evan: Let's start with the multi-agent orchestration.
The evaluation uses the Multi-Agent Orchestration Benchmark, or MAOB, which comprises 140 requests modeled on occupational tasks.
Each task is paired with a reference directed acyclic graph, or DAG, of four specialist sub-agents.
Ashley: The specialists cover research, coding, content creation, and on-call tasks.
The benchmark measures agreement between the host's proposed graph and a reviewed reference graph.
This allows us to evaluate specialist selection and dependency planning independently of downstream execution quality.
Evan: Exactly.
The metrics used include Node F1, Edge F1, Partial Order Accuracy, and Exact Match Rate.
Raven's orchestration performance is compared to two baseline harnesses—Claude Code and Hermes Agent—under two backbones, Qwen3.8-27B and DeepSeek-V4-Flash-0731.
Ashley: Raven consistently outperforms both baselines across all metrics and backbones.
For instance, with Qwen3.8-27B, Raven achieves a Node F1 of 0.923 compared to 0.776 for the strongest baseline.
Edge F1 also shows improvements, with Raven reaching 0.950 under the same backbone.
Evan: Raven's Exact Match Rate is 0.711 with Qwen3.8-27B, offering a significant 10.4 percentage-point improvement over the strongest baseline.
With DeepSeek-V4-Flash-0731, Raven achieves an Exact Match Rate of 0.867, a 10.5 percentage-point improvement over the strongest alternative.
Ashley: These metrics show Raven's strong capability in specialist selection and accurate dependency planning, significantly outperforming alternatives.
The results highlight the robust orchestration capabilities of Raven.
Evan: Next is harness self-evolution.
Raven builds on HarnessBank and evaluates the diagnosis, search, and screening procedures on tasks withheld from evolution.
Seven benchmarks are used, including Terminal-Bench-2, EvoAgentBench covering five domains, and AppWorld.
Ashley: The evolved harness consistently improves held-out Pass@1 across all benchmarks.
For example, harness evolution raises the score from 36.1% to 45.4% on Terminal-Bench-2 and from 41.3% to 56.7% on AppWorld.
The improvements demonstrate the effectiveness of harness adaptation and the transferability of gains across different tasks.
Evan: It's impressive how the evolved harness achieves such notable gains, confirming the utility of Raven's approach to modular harness evolution.
Ashley: Moving on to Raven-Research, which answers deep research questions using evidence from the live web.
Evaluated on DeepResearch Mixed, it combines questions from BrowseComp, FRAMES, Humanity’s Last Exam, and xBench-DeepSearch.
Evan: On DeepSeek-V4-Flash, Raven-Research achieves a pooled accuracy of 76.5%, outperforming all compared systems.
With Qwen3.6-35B-A3B and Qwen3.5-397B-A17B, it also leads with 56.3% and 59.3% accuracy, respectively.
Ashley: In addition to accuracy, Raven-Research balances inference costs well.
For example, on DeepSeek-V4-Flash, its mean cost per question is $0.0242, achieving higher accuracy at competitive costs.
Evan: Then we have Raven-Code, which focuses on repository-aware software engineering.
It resolves tasks from SWE-bench Pro and Verified, repository development from WorkBuddy Bench, and whole-repository migration from SWE-Refactor.
Ashley: Raven-Code achieves high scores across benchmarks.
On SWE-bench Pro, it resolves 15 more tasks than Claude Code with Qwen3.8-27B.
On SWE-Refactor, Raven-Code outperforms leaderboard entries by 9.5 points with DeepSeek-V4-Flash.
Its combined performance shows strong results on complex repository tasks and targeted bug fixing.
Evan: DataAgentBench further confirms Raven-Code's strengths, achieving the highest Pass@1 of 0.8762 among leaderboard entries for analytical database queries.
Ashley: Finally, Raven-Design handles visual design and inspection.
PresentBench overall scores show Raven-Design outperforming baselines in both tested backbones.
It scores 80.2 with Claude Opus 5 and 72.9 with GPT-5.6 Luna.
Evan: Raven-Design achieves consistent gains across visual-design benchmarks, Dashboard, SVG, and GDPval tasks.
It led in all settings, demonstrating its effectiveness in visual artifact generation and professional deliverables.
Ashley: And we see similar effectiveness in managing long-running jobs with Raven-Oncall.
For AI4AI, Raven-Oncall achieves better pretraining quality and lower costs under Claude Opus 5 and DeepSeek-V4.1-Flash.
It also solves more tasks with lower resource use in the AI4S scientific benchmark.
Evan: Across all experiments, Raven consistently outperforms comparable systems, demonstrating the utility of its orchestration, adaptation, research, coding, design, and sustained task management capabilities.
Ashley: This concludes the Experiment section of the paper review.
Evan: Now, let's delve into the Related Work section to see how Raven builds on and compares with existing methods in the field.
Ashley: The paper highlights four main areas of related work that Raven integrates and advances: multi-agent collaboration and orchestration, agent harness self-evolution, long-term memory for agents, and skill acquisition, retrieval, and evolution.
Evan: Starting with multi-agent collaboration, different frameworks have been developed to organize agents within a shared framework.
AutoGen and MetaGPT are notable examples that specialize through configurable message exchanges and role prompts.
Ashley: Yes, AutoGen allows developers to compose agents through messages and conversation patterns, while MetaGPT assigns roles for collaborative software development.
These frameworks support specialization through communication configurations but often lack an explicit typed graph that can be validated before execution.
Evan: Other systems like MacNet represent interaction with directed acyclic graphs, and Magentic-One maintains a task ledger, assigning agents dynamically based on changes during execution.
MasRouter, meanwhile, learns to route requests among large language models in multi-agent setups.
Ashley: This demonstrates that coordination gains depend heavily on the interaction between task structure and system architecture.
Kim et al.
found that increasing agent count alone doesn't predict collaboration performance.
Efficient coordination is essential, especially in complex task environments.
Evan: Raven's approach to multi-agent orchestration includes a Host Agent that delegates tasks through explicit executable graphs, ensuring dependencies are well-formed and validated before execution.
This separates planning from execution, which allows for better coordination across diverse agents.
Ashley: Moving on to agent harness self-evolution, previous methods like GEPA and ADAS focus on evolving prompts and agent designs.
GEPA evolves prompts through reflective feedback, while ADAS automates the design of agentic systems.
These approaches optimize specific parts of agent execution without providing a comprehensive mechanism for evolving the harness itself.
Evan: Open-ended evolution systems like the Darwin Gödel Machine and Growing Harness address broader aspects by evolving the agent's own code or its harness around a frozen model.
They work within the constraints of an offline optimizer or reusable scaffolding respectively.
Ashley: HarnessBank introduces a semantic gene bank approach, allowing modular harness evolution with gated verification.
This supports extensive evaluation and comparison across various harness modifications categorized by prompt, knowledge, runtime, and configuration edits.
Evan: Raven builds on these methods by integrating harness evolution directly into its multi-agent system.
The modular design allows changes to the harness without altering the underlying model.
The fixed evaluator and Evolver Agent ensure that only the most effective harnesses are retained.
Ashley: Next, we consider long-term memory for agents.
Memory systems range from context management and reflection models like Generative Agents and MemGPT to scalable systems like MemoryOS and MIRIX, which organize memory across different dimensions and substrates.
Evan: These systems illustrate varied approaches to managing memory.
Generative Agents simulate human behavior and reflect on episodes, whereas MIRIX includes procedural memories alongside other types in a cohesive system.
However, the difference between task artifacts and memory records often remains blurred in practice.
Ashley: Raven leverages EverOS to structure memory as discrete episodes, facts, and reusable skills, enabling efficient recall and update processes.
By separating artifact completion from memory recording, Raven ensures clear boundaries between what an agent needs immediately and what it retains for future use.
Evan: Lastly, skill acquisition, retrieval, and evolution are critical for effective agent performance.
Early systems like Voyager and ExpeL learn executable code and natural-language insights, respectively, while others like SkillCorpus curate extensive catalogs of public skills.
Ashley: SkillCorpus, for example, aggregates, deduplicates, and evaluates public skills, providing a curated database for agent retrieval.
SkillRouter enhances this retrieval with specialized models, ensuring appropriate skill routing at scale.
Evan: Memp and SkillWeaver represent the latest steps in procedural skill evolution, updating knowledge dynamically based on execution experiences.
They however often focus on within-task gains rather than a broader multi-agent system perspective.
Ashley: Raven's Skill Forge extends this by integrating local and EverOS-derived skills with SkillHub, enabling the agent to retrieve and utilize reusable procedures.
Modular updates from execution cases ensure skills remain relevant and improve over time.
Evan: In summary, Raven stands on the shoulders of previous work while forging new paths in multi-agent collaboration, harness evolution, agent memory, and skill utilization, making it a comprehensive solution for complex, long-horizon tasks.
This concludes the Related Work section of the paper.
Evan: As we near the end of today's episode, let's summarize the key contributions and takeaways from the paper.
Ashley: The paper presents Raven, The Harness of Harnesses, an innovative multi-agent ecosystem designed to tackle the complexity of harness development and the composition of specialized capabilities across domains.
Evan: A standout feature of Raven is its advanced multi-agent orchestration.
The system's Host Agent efficiently decomposes tasks, matches subtasks to specialized agents, and ensures execution dependencies are well-coordinated.
Raven's orchestration capabilities significantly outperform existing systems, as demonstrated by its superior Exact Match Rate on the MAOB benchmark.
Ashley: Another core contribution is harness self-evolution.
Building on the foundation of HarnessBank, Raven's modular approach allows for continuous improvement and effective adaptation of agent harnesses.
The rigorous evaluation and screening processes ensure only the most effective harness modifications are implemented.
Evan: Raven also excels in task performance across various domains.
Its specialized agents—Raven-Research, Raven-Code, Raven-Design, and Raven-Oncall—demonstrate remarkable efficiency and accuracy in their respective tasks, from deep research to repository-aware software engineering, visual design, and sustained operation tasks.
Ashley: Moreover, Raven's memory systems and Skill Forge facilitate procedural reuse and skill improvement.
By integrating local skills, EverOS-derived skills, and SkillHub, Raven ensures that agents leverage past experiences to inform and enhance future tasks.
Evan: In essence, Raven stands as a comprehensive solution, pushing the frontiers of composable agentic intelligence.
It combines effective task orchestration, modular harness evolution, and advanced memory and skill management to tackle complex, long-term tasks successfully.
Ashley: Thank you for joining us on this episode of Daily Paper Cast.
We hope you found our discussion on Raven insightful and thought-provoking.
Evan: Be sure to tune in again for more deep dives into the latest advancements in AI research and technology.
Until next time, stay curious and keep exploring.
Goodbye!