The Harness

Mozilla's report prices the commoditization gap

Show Notes

Mozilla's new open-source AI report puts a number on the commoditization paradox: open models now handle a third of production tokens but capture only four percent of the revenue. Databricks answers with a bet of its own, raising at a $188 billion valuation on an explicit harness-not-model strategy while Isomorphic Labs pushes AI substitution past diagnosis into drug design. Meanwhile a beloved coding benchmark quietly died as a signal, and Kaiser nurses turn AI workplace surveillance into a live labor fight.

What is The Harness ?

A daily summary of what is interesting and happening in the AI industry, with a focus on what this means for people building harness experiences that are used.

Good morning, it's Saturday, July eighteenth.

In today's briefing, we see open-weight models now handling a third of production tokens but capturing only four percent of the revenue, a major investment in orchestration infrastructure from Databricks, and AI moving past diagnosis into drug design.

In the harness, tools and orchestration world;

Databricks has raised approximately three billion dollars at a one hundred eighty-eight billion dollar valuation on a pitch that's explicitly not "best model." The round, led by Coatue and up from one hundred thirty-four billion in February, centers on Lakebase, an agent-native database; Unity, an AI gateway; and Omnigent, a meta-harness for coordinating multiple agents. While Databricks champions cheaper open models like GLM five point two for coding work, the real thesis is that durable value sits in the orchestration and data layer around a commoditized model, not the model itself. For product teams building full-stack agentic systems, expect the architectural discussions to shift from "which model" to "how does the routing and memory layer handle state," because frontier labs have commoditized the core reasoning piece and the moat has moved to orchestration.

Simon Willison's analysis on Hacker News addresses his own pelican-riding-a-bicycle benchmark: it no longer works as a signal. Models like GLM five point two now beat Fable five and GPT five point six on it without being frontier class, and the test never measured what matters now, which is agentic tool calling. The deeper issue is that benchmark gaming isn't just a vendor problem; community-built tests get contaminated once enough people optimize against them. For procurement and evaluation processes, this means moving past generation-quality tests toward task-grounded, tool-use measurements, because the optimization arms race has reached even well-intentioned community signals.

In other lab news today, Isomorphic Labs unveiled a Drug Design Engine that more than doubles AlphaFold three's accuracy on protein-ligand prediction, beats gold-standard physics methods on binding affinity, and discovers novel binding pockets from sequence alone. Demis Hassabis expects AI-designed drugs in clinical trials by year-end. This is a step past diagnostic-substitution stories where AI reads existing medical images; AI is now proposing the molecule itself. For teams in regulated industries using AI for discovery work, the verification burden shifts upstream, because the human check becomes trial design and regulatory sign-off, a much harder comprehension problem than interpreting an existing image.

In other news, Mozilla's inaugural State of Open Source AI report, led by CTO Raffi Krikorian, quantifies what this briefing has tracked for weeks: the gap to closed frontier models has narrowed to three percent, inference pricing has dropped fifty times in three years, and open weights now handle roughly a third of production tokens. Here's the problem: open models capture only four percent of the revenue. Usage share and value capture have completely decoupled. For teams building products on open weights, this means the model itself cannot be your moat; something else must be, because the market has priced model switching costs essentially to zero.

Kaiser nurses on advice and triage lines report that productivity software and a paused-but-potentially-returning AI tool that scores empathy in their voices are compressing their time with patients in crisis. The California Nurses Association opens new contract talks this month with AI central to the negotiation, following strikes in March and last fall. This isn't AI doing clinical work; it's AI as a real-time management layer shaping how labor itself operates. For product teams shipping automation into frontline environments, expect labor agreements to shape your design constraints, because union negotiations and regulatory frameworks now govern how these tools operate.

That's the briefing. Have a great day.