The Harness

IPO theatrics meet infrastructure churn

Show Notes

Anthropic pitches IPO investors a $30 trillion addressable-market forecast and $190-200B in 2028 revenue, landing the same week OpenAI loses its 13th executive of the year and gets independent, partially-caveated verification of its new Jalapeño inference chip. Elsewhere, three separate research threads converge on agent harness design mattering more than model choice, the EPA moves to strip public comment from data-center air permits just as investors price backlash risk into Anthropic's prospectus, and OpenAI dismantles a fabricated Russian-linked think tank built on ChatGPT-drafted content. Plus: Alibaba's programmable-memory agent architecture, Figure's billion-dollar robotics data bet, and OpenAI and Anthropic converging on the same enterprise agent-governance playbook.

What is The Harness ?

A daily summary of what is interesting and happening in the AI industry, with a focus on what this means for people building harness experiences that are used.

Good morning, it's Wednesday, August twenty-sixth.

In today's briefing, independent testing narrows the claims behind OpenAI's new inference chip, Anthropic pitches IPO investors a thirty-trillion-dollar market forecast the same week OpenAI loses another executive, and three separate research threads agree that agent harness design now matters more than model choice.

First up, today in the big model news.

OpenAI
SemiAnalysis independently benchmarked OpenAI's new Jalapeño inference chip and largely backed the headline claim: it leads every chip tested on tokens per megawatt, with strong single-concurrency performance on DeepSeek R1 and Kimi-K2.5. The verification comes with real caveats: no full benchmark suite, no long multi-turn agent testing, a comparison against Nvidia's Blackwell that the assessor itself calls unfair since the real rival is the still-unreleased Vera Rubin chip, and only the earliest chip revision under test. OpenAI's own Astra and Codex models reportedly helped write the chip's low-level kernels, roughly one and a half to two times faster than the existing code, in about two months. That narrows the compute story more than it confirms it: efficiency gains look real for short-context inference, still unproven for the agentic workloads that actually drive production cost.

Frontier economics, not model capability, keeps driving the headlines out of the top labs. OpenAI lost its data center chief, Chris Malone, its thirteenth executive departure this year, just as it tries to lock in roughly six hundred billion dollars of compute commitments through twenty thirty and close its enterprise sales gap with Anthropic.

Anthropic
Anthropic is telling its own IPO investors that its addressable market tops thirty trillion dollars, larger than SpaceX's benchmark pitch and close to the size of the entire US economy. It's projecting revenue of one hundred ninety to two hundred billion dollars by twenty twenty eight, building on a second quarter that already doubled to eleven point six billion dollars and delivered the company's first operating profit. And this part is contested: outside analysts and the Financial Times question whether that growth reflects more customer spending, or just cheaper performance sold at a lower price. The imminent prospectus is the real test: it will show whether the thirty-trillion-dollar framing survives underwriters rather than press coverage.

In the harness, tools and orchestration world;

Microsoft-adjacent researchers published AutoSaddler, a new agent harness posting double-digit gains on GAIA2, SWE-Bench Pro, and Terminal-Bench 2.0 over baseline harnesses running identical models. It's the third independent research thread to reach the same conclusion, after GLM-5.2's harness-only jump from twenty-three to fifty-two percent pass rate and Nvidia's new skill-lift metric: harness configuration now swings benchmark scores more than model choice does. AutoSaddler's own numbers aren't independently verified yet, but the researchers are already proposing a Harness Card, a model-card equivalent for scaffolding, since comparisons that don't disclose it are close to meaningless. A model's leaderboard score means little now without knowing which harness produced it.

Alibaba's latest research treats agent memory as programmable state: append-only event logs plus a persistent Python kernel, instead of compressed chat history, and it reports strong long-session retention on Qwen3.8-Max. The score is self-reported and unverified, but the architecture bet, treating memory as queryable state rather than squeezed-down history, is the more interesting signal here.

OpenAI shipped an Admin plugin adding workspace management and permissions controls to ChatGPT Work and Codex, the same week it disclosed Codex crossed twenty million active users and needed org-level spend controls after heavy users burned through budgets in days. Anthropic built the equivalent layer days earlier: enterprise-managed authentication for its MCP connectors across Asana, Atlassian, Slack, Notion and other tools, following the MCP maintainers' own plan to replace shared API keys with per-agent identity. Both labs are converging on the same problem, central IT control over what agents can touch, from opposite directions. The contested ground in agent systems has shifted to the space between agents and the tools they touch, not inside any one model.

In other news…

OpenAI banned a VPN-routed cluster of Russia-linked accounts that used ChatGPT to draft content for the International Burke Institute, a fabricated think tank with invented contributors, including false bylines for Francis Fukuyama and Noam Chomsky, and plagiarized or misattributed thirty-four of its thirty-six articles. Its core product, a so-called sovereignty index rating countries favorably toward Russia and critically toward the US, France, Germany and the EU, was built with other tooling; OpenAI's models handled only the surrounding content. OpenAI says reach was small despite the elaborate setup, following its June disruptions of separate campaigns, including a Chinese-linked one, targeting US AI and data center policy debates. The shift from one-off jailbreak posts to persistent fabricated institutions means source verification now has to cover entire fake organizations, not just individual articles.

Figure is running the LLM playbook on robots: its new Index dataset has already pulled in sixteen million video uploads and paid contributors fifteen million dollars, and the company is committing a billion dollars over the next year to buy data scale before betting on new architecture. A rival humanoid robotics company raised nearly a billion dollars on the identical thesis. Data scale, not a new model architecture, is where humanoid robotics money is going right now.

On the regulatory front, the EPA has proposed ending the fifty-year-old requirement that states give thirty-day public notice before issuing Clean Air Act minor-source permits, the category covering data center diesel generators, making public participation optional and left to state and local discretion, many already tied to data center operators by business agreements. The agency frames it as cutting administrative burden for economic development, a position that contradicts its own administrator's March transparency memo, and groups including the Sierra Club and the Center for Biological Diversity have formally opposed it. It arrives as a federal counterweight to state and tribal restrictions on data center buildout, from Texas's grid operator pausing new connections to the Cherokee Nation's land ban, the same week investors are separately asking Anthropic to price public backlash risk into its own IPO prospectus. Lowering the disclosure bar as that backlash keeps growing turns permitting timelines into a real variable for anyone building on this infrastructure, not just a compliance footnote.

That's the briefing. Have a great day, and don't forget to subscribe.