Maya: One engineering team let Cursor's coding agent merge roughly two thousand five hundred pull requests into production in a single month. And they frame that number as a side effect, not the goal. James: Right, the real project was building an environment they could trust enough to stop hand-holding every agent. The walkthrough was prepared for Cursor's Compile conference in London, then posted free on X. Maya: The problem that started it was performance regressions. New pull requests arrived faster than one engineer could hand-check each for slowdowns, so his manual trace routine got written down as instructions an agent could follow. James: The mechanism is four layers. Control Glass lets agents run the app themselves through Chrome's DevTools Protocol to capture real performance traces, rather than trusting that code compiles. A feature map stores where each feature lives so an agent can reproduce a vague bug. Maya: And their blunt rule ties it together. Any time a person corrects the same mistake twice, that correction gets folded into a lint rule, a test or a dependency boundary, so the next agent run starts from an already-improved environment. James: There's a related workaround worth explaining. Grok Bot, xAI's assistant, has no setting to switch which model answers, so builders forward tougher jobs to Cursor Cloud, where a different model can take over. Maya: The clearest demo happened mid-bike-ride. A builder found a broken GPX file, spoke through open-ear headphones asking a Cursor Cloud agent running Fable five point one to fix it and re-upload to Ride with GPS. James: And the fix had synced to a Garmin bike computer by the time the ride ended. That same routing covers non-code work too, like a complicated letter needing heavy evidentiary support, plus plan requests and video tasks. Maya: One builder even downgraded a SuperGrok subscription to standard while keeping Cursor, hoping the two products fold into a single subscription within a few weeks. Until then, the full benefit means paying for both. James: Cursor's Plan Mode takes a different approach to trust. It adds a pause between your prompt and the code, first asking clarifying questions to fill gaps in its understanding of your project. Maya: Then it writes a plain-text roadmap laying out which files it plans to change and in what order. You can edit that directly, deleting steps built on wrong assumptions or adding ones it missed. James: Once approved, that roadmap becomes a strict boundary. Cursor executes the whole plan in a single coordinated pass across files, rather than improvising as it goes. In the test case, the login system came out built correctly on the first attempt. Maya: And the pitch is for changes where a fix in one place silently breaks another. The video's example is a login change that could interfere with an interconnected checkout system. James: That planning discipline connects to a developer who says he ships software from a beach in Tahiti, running Cursor's Projects feature alongside a tool called PStack. He calls it the best mobile setup he's found for handing coding tasks off to agents. Maya: He says the setup builds what feels like infinite memory of how he ships. He taught it his QA flow and his devops release process, and it applies those habits later without him repeating instructions. James: The other piece is asynchronous tracking. He gets updates on everything the agents finished overnight, so he can review and approve work without having been at a keyboard while it happened. Maya: He names one limit himself, James. He'd like the setup connected to Grok Bot, but says that integration doesn't exist yet. It's a wish, not a current feature. James: Last one is a small toggle worth knowing. Cursor three point zero's default is agent-centric, and it only recaps changes after the fact, hiding which files the agent is actually editing while it works. The fix is a button labeled IDE in the top right corner. Maya: Clicking it restores the file tree on the left, so edits show up live as the agent works. It matters because agents now run much longer unsupervised turns, and the demo showed Grok four point six running its own internal QA before presenting a result. Maya: That's your look inside Cursor HQ for now. Try that IDE toggle while you're still learning the agent, and we'll see you back here soon.