The Harness

Anthropic and OpenAI lobby against the openness they defend

Show Notes

Nvidia rallies 35 companies into an Open Secure AI Alliance the same day Anthropic denies wanting an open-weights ban, hours after the New York Times reports Anthropic and OpenAI have been privately lobbying Washington to restrict Chinese open models. Claude Opus 5 tops day-one leaderboards but barely improves on a benchmark built to catch code-quality decay, and a $500 fine-tuned open model outperforms five frontier options on a real e-commerce task at a fraction of the cost. Kimi K3's open weights hit a hardware wall requiring Blackwell-class GPUs, while Google and Microsoft race to become the default AI layer inside a US national science initiative.

What is The Harness ?

A daily summary of what is interesting and happening in the AI industry, with a focus on what this means for people building harness experiences that are used.

Good morning, it's Tuesday, July twenty-eighth.

In today's briefing, Nvidia rallies dozens of companies into a security alliance the same day Anthropic and OpenAI are revealed lobbying against the open models they publicly defend, Claude Opus 5 tops leaderboards but barely moves a benchmark built to catch code quality decay, and a five hundred dollar fine-tune beats five frontier models on a real task.

First up - Today in the big model news;

Google
Google DeepMind committed forty million dollars in AI tokens and cloud credits to the Department of Energy's Genesis Mission, the national push to double the pace of scientific discovery within a decade, giving researchers at all seventeen DOE national labs access to Gemini and to AlphaEvolve, its algorithm discovery agent. Microsoft matched days earlier with sixty million dollars of its own in-kind commitment.
Neither company is writing a check. They're donating compute and model access to become the default AI layer inside a government science initiative before the other locks in the relationship. The signal to watch is which lab's tools researchers actually adopt once the credits go live.

Anthropic
Claude Opus 5 topped Arena's Frontend Code and Text leaderboards on its first day out. But an independent benchmark called SlopCodeBench, built specifically to measure code quality decay across nearly two hundred checkpoints rather than one shot pass rates, put Opus 5 at a twenty four percent strict pass rate, just seven points above Opus 4.6's mark from four months ago, with structural erosion still showing up in more than three quarters of trajectories. That lines up with early developer feedback describing Opus 5 as overcomplicated and slow to stop once a task is done.
A model can top a headline leaderboard and barely move a benchmark built to catch what headline leaderboards miss. If you're routing production coding work off launch week rankings, the decay number is the more decision relevant one, and it's the one Anthropic hasn't addressed yet.

Moonshot AI
Kimi K3's open weights, released with a hundred and four billion active parameters in MXFP4 format, hit a hardware wall. A cluster of eight A100s can't fit the weights without multi node sharding, eight H200s still need two nodes, and only eight B300s, roughly five hundred thousand dollars for early testing, comfortably run the model with room left for long context memory.
Open license and runnable are turning out to be two different claims, and a day after the weights shipped, that gate is still holding. Budgeting for K3 means pricing the hardware alongside the license from day one.

In local model developments, a five hundred dollar fine-tune just outperformed five frontier models on a real task. FermiSense trained a small open-weight Qwen model on one e-commerce catalog review workflow and beat five frontier options, including the best prompted configuration, scoring nearly thirteen points higher on a quality index, at fifty cents per thousand listings reviewed against thirty four dollars for the strongest frontier model. The telling detail: all five frontier models plateaued within a tenth of a point of each other no matter how they were prompted, while the trained specialist cleared that ceiling entirely.
That's a small model trained on owned task data beating every frontier option combined, at roughly a seventieth of the cost, not a harness wrapped around a rented frontier model. For well defined, high volume tasks, the first procurement question stops being which frontier API to use.

In the harness, tools and orchestration world, a clean benchmarking post ran the identical DeepSeek V4 Flash model across three coding harnesses, Pi, OpenCode, and Claude Code, and found the same model finishing an identical task in two point one minutes, three point one minutes, and eight minutes respectively, purely a function of prompt size and tool calling structure, not model capability. It lands the same week Anthropic is publicly defending its own system prompt cuts as a deliberate simplification bet.
Harness choice is moving outcomes as much as model choice now. Benchmark the model and harness pair together, never the model alone.

On the AI policy front today, Nvidia rallied more than thirty five companies, including Microsoft, IBM, Cisco, Dell, Salesforce, SAP, and CrowdStrike alongside Hugging Face, LangChain, Nous Research, and Unsloth, into a new Open Secure AI Alliance built around Jensen Huang's claim that attackers already have strong AI, so defenders need both open and closed frontier models. The launch pointedly referenced the OpenAI and Hugging Face sandbox breach from two weeks ago: a closed model blocked Hugging Face's own forensics team mid response, while an open weight model helped contain the intrusion.
Anthropic, notably not a founding member, used the same day to publish a clarification that it has never advocated for a ban on open weight models, and Dario Amodei dismissed open source as a red herring. Hours earlier, the New York Times had reported that Anthropic and OpenAI have been quietly pushing Treasury Secretary Bessent and White House tech adviser Kratsios to restrict Chinese open weight models. OpenAI declined to join Nvidia's alliance at all, which triggered internal employee pushback, and only signed a separate twenty five firm open weights coalition letter after its absence became a story online. Anthropic still hasn't signed.
The labs that have spent two months supplying the safety rationale for state restriction are the same ones shown privately lobbying for it while publicly disclaiming it, and the infrastructure vendors who would profit from open weight demand are the ones defending it in public. Watch whether Commerce takes the narrower, case by case route insiders describe as favored: that looks built for the private lobbying position, not the public statements.

That's the briefing. Have a great day.