Vercel's vision made one production agent easier to ship. The second agent reveals the hidden cost: shared context, permissions, handoffs, approvals, and proof.
Episode 39: I Built More Agents. The Work Got Harder.
Vercel's vision made one production agent easier to ship. The second agent reveals the hidden cost: shared context, permissions, handoffs, approvals, and proof.
Hosts: Luis and Carina · Runtime: 11.3 minutes
What product managers can actually build with AI today—and where it still breaks.
LUIS: Most agent demos end at the satisfying moment: the agent returns an answer.
CARINA: And your work usually starts there.
LUIS: Exactly. I ask one agent to research a market. It returns a useful brief. I pass that brief to another agent and ask for a plan. The second agent asks for the sources again. The third needs to know which parts are facts and which parts are guesses.
CARINA: Then somebody asks the question no demo wants to answer: who approved the next step?
LUIS: Nothing looks broken. Each agent did something reasonable. I still have to copy the context, check the sources, reconcile two slightly different answers, remember which decision is current, and approve the same workflow from another screen.
CARINA: So the agents are doing the work, and you are carrying the story between them.
LUIS: That is the pain. The system gets faster at producing outputs, but slower for a human to understand and control.
CARINA: Welcome back to Contextually Aware. I’m Carina.
LUIS: And I’m Luis. Today we’re talking about the moment an agent project stops feeling like one useful tool and starts feeling like a small organization nobody has managed to organize.
CARINA: The title is “I Built More Agents. The Work Got Harder.” That sounds like a confession.
LUIS: It is. I was counting agents, tools, models, and skills. I was not counting the time I spent making sure they understood one another.
CARINA: Give us the scale of the problem.
LUIS: In a late-July inventory of my own setup, I counted 13 harnesses, which are the software shells that run agents; 12 model IDs; 48 agent definitions; 390 skills; and 27 MCP servers. MCP is an open standard for connecting agents to tools and data.
CARINA: That is a lot of machinery for one human reviewer.
LUIS: It is a dated snapshot, not a benchmark. The point is what happens when execution grows faster than coordination. I had built the doing faster than I had built the coordinating.
CARINA: Why is the first agent usually fine? You can give it a task, inspect the answer, and improve it.
LUIS: The feedback loop is short. Give it a task, watch it work, check the result, fix the instructions, run it again.
CARINA: And then the second agent arrives.
LUIS: The second agent changes the shape of the work. Now someone has to decide which agent gets the task, what information it receives, what it can do, and what the next agent can trust.
CARINA: That sounds less like a model problem and more like an operations problem.
LUIS: That is how it felt. Every agent had its own private idea of the context, the tools it could use, where the work stood, and what counted as proof.
CARINA: Walk me through a failure that people will recognize.
LUIS: One agent researches a market. Another turns the research into a brief. A third writes a draft. A reviewer checks it.
CARINA: And every unanswered question became your job.
LUIS: Right. I was the router, the shared memory, the permission check, and the audit trail.
CARINA: What does a handoff look like when there is no shared system?
LUIS: Usually one of four things. A short summary that drops the qualification that mattered. A full transcript that keeps everything but makes the scope hard to see. A human translation step, where I restate one agent’s work for another. Or a silent assumption that one agent’s private notes are shared truth.
CARINA: The last one sounds dangerous because the workflow can look successful.
LUIS: Exactly. A stale map of buyers gets treated as current. A draft uses a source the next agent cannot check. A tool permission follows the agent instead of the task. The workflow finishes, but nobody can explain why the result is safe to use.
CARINA: Was there a specific moment when this became clear to you?
LUIS: I watched Guillermo Rauch’s Vercel talk, [The AI Agent Every Company is About to Build](https://youtu.be/HQXi4snP36I), and then made a presentation about it.
CARINA: What did you take from the talk?
LUIS: The basic idea is that software is gaining a new kind of worker: the agent. An agent is a program that uses an AI model to do real work on its own. It can write code, ship it, check the result, and continue after a person walks away from the keyboard.
CARINA: That changes the infrastructure around software.
LUIS: It has to. A production agent needs somewhere to work, a way to stop, a way to recover after an interruption, boundaries around what it can touch, and a record of what happened. It also needs a human approval step when the risk goes up.
CARINA: So the talk helped with the first layer: making one agent real.
LUIS: Yes. It also made me notice that AI models are not the only systems with context windows. The operator has one too.
CARINA: This is where Vercel’s eve enters the story.
LUIS: Vercel describes eve as a framework that grew from teams rebuilding the same production plumbing for internal agents.
CARINA: Put that in ordinary language.
LUIS: Eve gives an agent a durable place to work. Sessions and workflows can pause and resume without losing their place. A sandbox gives the agent a walled-off environment. Approvals can wait for a human. Subagents can help with smaller jobs. Traces record what happened. Quality checks test behavior. Git tracks changes.
CARINA: That sounds like a real software system, not a prompt in a chat window.
LUIS: That is the value. Eve handles the production shape of one agent.
CARINA: And what remains when you add the second agent?
LUIS: A directory can describe one agent. A runtime can resume one workflow. Neither automatically tells the next agent which decision to inherit, which context to leave out, or what authority was handed over.
CARINA: Do the other agent frameworks solve that missing layer?
LUIS: They solve important pieces. Google’s Agent Development Kit, or ADK, separates the current session from working state and longer-lived memory. That helps distinguish what happened in this run from what the organization has learned over time. ADK also documents routing and evaluation.
CARINA: What about OpenAI’s Agents SDK?
LUIS: It makes two coordination choices explicit. A manager agent can call specialists as tools. Or one agent can hand control to a specialist. That changes who owns the conversation and the next decision. The SDK also documents guardrails, which are checks on what an agent takes in and produces.
CARINA: And LangGraph?
LUIS: LangGraph is a lower-level option for custom workflows. It supports saved state, durable execution, streaming output, and points where a human can step in.
CARINA: MCP and A2A sound more like connection standards.
LUIS: They are. MCP gives agents a common way to expose tools, resources, and prompts. A2A, or Agent2Agent, describes communication between independent agents. These standards help agents discover capabilities and exchange tasks.
CARINA: But they do not decide who owns a task or who approves a risky action.
LUIS: Exactly. That decision belongs to the company operating the system.
CARINA: So what is your recommendation?
LUIS: I call it the Agent Workplane. This is my synthesis and recommendation, not an existing industry standard.
CARINA: The runtime can change underneath that contract.
LUIS: Exactly. Eve, ADK, the OpenAI Agents SDK, LangGraph, or a future runtime can sit underneath as a replaceable engine.
CARINA: How do you roll it out without creating another giant platform project?
LUIS: Three sequences: Establish, Operate, Compound.
CARINA: Start with Establish.
LUIS: Write the charter. Who owns the agent? What is it for? What is it forbidden to do? What data belongs in its context? How will we measure whether it helps? Then give it one useful job with a visible baseline. Keep the context scoped and the tools limited. Make the first version read-only wherever possible.
CARINA: Operate sounds like the coordination step.
LUIS: It is. One front door. One router. A few bounded workers with narrow skills and clear contracts. A research worker can read approved sources. A content worker can draft. A build worker can run checks. None of them inherits every permission the company has.
CARINA: And approval becomes part of the task, not a side conversation.
LUIS: Right. The agent pauses. The human sees the exact action and the evidence. The approval is recorded. The run resumes from the same checkpoint.
CARINA: Then Compound is where independence is earned.
LUIS: Yes. Events and schedules can start work. Feedback and evaluations can find repeated failures. The system can propose a new skill or route, but a human approves the change before it becomes part of the operating contract.
CARINA: What would you actually ship first?
LUIS: One front door, one router, and three workers with clear limits. Read-only context. Human approval before anything writes data or becomes visible outside the company.
CARINA: Give us a concrete test.
LUIS: Give the system 10 to 20 named accounts. Ask it to produce five account dossiers. Each dossier includes facts, likely buyers, source links, unknowns, and a confidence note.
CARINA: What do you measure?
LUIS: Time against the old process, cost, response time, source errors, recovery time, and any actions that happened without review. For the first version, that last number should be zero when there are outside consequences.
CARINA: That answers a better question than “Which framework should we choose?”
LUIS: The better question is: can the system keep the contract while doing real work?
CARINA: Let me see if I have the takeaway. The first agent removes execution work. The second agent exposes the coordination work that was hiding underneath.
LUIS: That is exactly it.
CARINA: So before adding another agent, define who owns the task, what context travels with it, what authority it receives, who approves the risky step, and what evidence gets saved.
LUIS: And make sure the business can keep those things if the model or runtime changes.
CARINA: I’m Carina.
LUIS: I’m Luis. This has been Contextually Aware.
CARINA: When did adding the next agent create more review work than useful output for you? Send us the failure mode. The best examples usually start where the demo ends.
LUIS: The next agent I add will have to earn its place in the workplane.