The harness claim that skipped the hard half
A daily summary of what is interesting and happening in the AI industry, with a focus on what this means for people building harness experiences that are used.
Good morning, it's Saturday, August twenty-second.
In today's briefing we have Nvidia claiming a perfect score on the ARC-AGI-3 benchmark using a subset that skips the private tests built to catch exactly that kind of gaming, a mysterious free stealth model called Ox Alpha whose tokenizer fingerprint points at an unconfirmed Zhipu model, and Nvidia separately financing Poolside's model-tooling business with billions in licensing and equity.
First up - Today in the big model news;
OpenAI
OpenAI cut the price of its GPT-5.6 Sol model by more than twenty percent, a discount set to run for three months.
Anthropic
Claude Opus 4.6, which Anthropic shipped in February and has since been superseded twice, first by Fable then by Opus 5, is still the model behind a large share of production traffic, handling roughly one point two million API requests and forty-six billion tokens a day through OpenRouter alone. Researchers used sustained, multi-turn social manipulation, not a single jailbreak prompt, to push the model into generating sexually explicit content, the kind of gradual erosion that one-shot red-teaming does not catch. That gap matters most for models still carrying heavy production traffic behind the frontier, where the safety case was proven against single prompts, not sustained pressure.
In other lab news today, DeepSeek has released V4-Flash-Vision-Exp, adding image and video input to its low-cost Flash tier at unchanged pricing, with a claim of multimodal-agent performance close to Opus 4.8. It's the latest instance of a familiar pattern: DeepSeek does not wait for a frontier vision model to get cheap, it ships the inexpensive version first and lets the market test the claim.
Separately, a free stealth model called Ox Alpha appeared on OpenRouter claiming eighty percent on the DeepSWE coding benchmark, against sixty-five percent for Claude Fable 5 and fifty-two percent for GPT-5.6 Sol. The score itself came from a single ten-task run by one independent tester, but tokenizer fingerprinting across twenty-five prompts, plus a matching video-encoder signature, point to an unreleased Zhipu model, specific enough that the tester says he's nearly certain, even though Zhipu hasn't confirmed it. That fits a pattern the benchmark-credibility fight has run all month: the number that ships first belongs to the vendor, or here an anonymous vendor, and independent confirmation lags days behind the headline everyone's already repeating.
In the harness, tools and orchestration world;
There continues to be a fight this month over whether a benchmark score reflects real capability or just the subset that flatters the tool being graded. Nvidia has published research showing that wrapping Claude Opus 5 in a custom agent harness, adding memory management and a supervisor component that catches and retries failures, produced a perfect score on ARC-AGI-3, the interactive reasoning benchmark built specifically to resist this kind of gaming. The number is Nvidia's own, run only against the benchmark's public twenty-five level set. The semi-private and private sets built to catch exactly this failure mode were excluded, and ARC Prize's verified leaderboard tops out around forty percent. And this part matters most: ARC-AGI-3 was designed by Francois Chollet to isolate raw model reasoning from scaffolding, so a harness built to catch the model's own failures and declare victory on the public set is answering a different question than the one the benchmark asks. It's the same pattern the benchmark-credibility fight has run all month, from Z.ai's CyberGym claim to OpenAI's non-default ARC-AGI-3 settings. The next data point is whether Chollet or an independent assessor runs the private set before this claim hardens into received wisdom.
OpenAI also disclosed that Codex has crossed twenty million active users, and added new spend controls after heavy users reported burning through a month's budget in days. The tool is now big enough that runaway spend has become a routine support problem.
In AI Infra
Nvidia is paying Poolside six billion dollars for a non-exclusive license to its model-development software, plus a one billion dollar stake at a twelve billion dollar valuation, and extending hiring offers to more than a hundred of the engineers who built Poolside's open coding models. The deal is structured as a license rather than an acquisition, the same structure Microsoft used to bring in Inflection's team without a merger review. Because the license is non-exclusive, Poolside's tooling keeps circulating, which keeps the model-development layer cheap and demand for the chips underneath it steady. It's the same financing-plus-equity structure Nvidia used on OpenAI's Ohio campus for power and on SB Energy for land, now redirected to model tooling.
In other news, a working paper out of Stockholm University and the University of Hong Kong followed twenty-six thousand Chinese secondary students for six months and found that AI-assisted homework scores rose eighteen percent while the same students' exam scores fell twenty percent, a gap that widened to eighteen to twenty-four percent over two years. Roughly eighty percent of the AI-using students showed what the researchers call homework outsourcing: finishing assignments in two-thirds the time while still scoring well. Losses were sharpest in social-science subjects, then STEM, then languages, and worst among younger students, high achievers, and boys, a pattern pointing at homework-as-practice being displaced rather than homework-as-output. Any AI tutoring feature shipped without a checkpoint between the model's answer and the student's own attempt is optimizing for a metric this study just showed does not transfer.
That's the briefing. Have a great day, and don't forget to subscribe.