The Harness

Frontier pricing power keeps eroding

Show Notes

Enterprise spending data shows Anthropic's flagship model capturing barely a tenth of its own customers' AI budget as cheap and open-weight alternatives pull share, forcing a price cut with the new Opus 5. An anonymous stealth model called Ox Alpha shows how fast hype outruns verification, while a proof-of-concept backdoor exposes a fresh supply-chain risk baked into coding-agent harnesses. Flock Safety's surveillance backlash and a copyright-law roundup round out a day where AI's legal and public-trust perimeter keeps tightening on multiple fronts at once.

What is The Harness ?

A daily summary of what is interesting and happening in the AI industry, with a focus on what this means for people building harness experiences that are used.

Good morning, it's Monday, August twenty-fourth.

In today's briefing we see Anthropic cutting the price of its new Opus 5 as cheap and open-weight alternatives eat into enterprise spending, an anonymous stealth model called Ox Alpha racking up hundreds of thousands of users before anyone could verify its benchmark claims, and a proof-of-concept backdoor showing how a coding agent's own harness can hand a hidden trigger to a malicious open-weight model.

Today in the big model news;

Anthropic
New data from Ramp and the Financial Times, covering seventy thousand companies, shows Anthropic's flagship model, Fable 5, has never captured more than about eleven percent of its own customers' Anthropic spending. DeepSeek prices under one dollar per million tokens, against Fable 5's ten to fifty dollar range. Open source model share on Vercel climbed from twenty-eight percent to sixty-two percent in two months. Anthropic answered not by defending its premium price but by cutting it: the new Opus 5 launched at five dollars for input and twenty-five dollars for output per million tokens. That concession landed three days after Anthropic's own IPO prospectus reportedly listed public backlash as a risk factor. Enterprise buyers are already routing spend to DeepSeek and open source models at scale, and Anthropic just followed the price down instead of holding the line. A flagship model's price is now reacting to commodity competition, not setting the market.

In other lab news today, a free, anonymous model called Ox Alpha appeared on OpenRouter and quickly became OpenCode's second most used model, racking up roughly sixteen trillion tokens and two hundred twenty one thousand users. An independent researcher says they're ninety nine percent certain it's an unreleased variant of Z.ai's GLM-5.3. Artificial Analysis lists no verified score for it, and the viral numbers don't hold up: an early ten-task sample put it near eighty percent on DeepSWE, but a run across a hundred and thirteen tasks found roughly sixty three percent. That gap between the early sample and the full run is why an unverified capability claim shouldn't drive production adoption. Treat a stealth model's viral numbers as marketing until an independent assessor publishes a real score.

In local model developments, a proof of concept trained a sleeper-agent backdoor into the open-weight Qwen 3.5 2B model that behaves normally until a preset date, then fires a malicious command instead. It worked on about eighty seven percent of in-distribution prompts and ninety percent of held-out prompts, with no misfires elsewhere. It specifically targets coding-agent harnesses, since tools like OpenCode and OpenAI's open-source Codex harness auto-inject the current date into every turn, handing the backdoor a built-in trigger with no external channel needed. It lands alongside a separate attack on the Rust package registry that hit packages with two hundred forty five million downloads combined. Downloading an unvetted open-weight checkpoint now carries the same supply-chain risk as pulling an unreviewed package. Audit what context your harness silently hands a model before trusting it with write access.

On the trust and legal front, Flock Safety CEO Garrett Langley responded to a Washington Post investigation that found at least forty six to fifty police officers misused the company's license-plate-reader network, including to stalk former partners, by arguing on Fox News that prioritizing only privacy or only safety is prioritizing the wrong thing, and calling for compromise. The backlash didn't slow after that: eighteen cities canceled Flock contracts in a single month, with Sanders proposing an outright ban and Republican Representative Burchett pushing to cut federal funding. The company's own defense concedes the tradeoff it wants regulators to avoid choosing between. Cities are pulling contracts faster than any federal rule could, so public trust, once broken, is being enforced locally instead of in Washington.

Courts have already settled that training on books counts as transformative fair use, in Bartz versus Anthropic, but how a training corpus was acquired is judged separately. Anthropic's one and a half billion dollar settlement, the largest copyright payout in U.S. history, paid for building a library out of pirated copies, not for training on them. Bartz versus Meta went Meta's way, but the ruling flagged an untested market-dilution theory as a stronger argument for future cases. Concord Music versus Anthropic is now testing whether Claude reproducing copyrighted lyrics on demand breaks fair use even when the training itself was lawful. The legal exposure for a product built on scraped text has moved from whether training is legal to how the corpus was sourced and what the model outputs on demand. Clean provenance and output-level filtering are becoming the real compliance requirement for any product built on scraped text.

That's the briefing. Have a great day, and don't forget to subscribe.