The Harness

OpenAI halts its own biggest training run over an unconfirmed threshold

Show Notes

OpenAI paused its largest frontier RL training run after Astra showed signs of crossing its own "Critical" cybersecurity threshold, tightening internal safeguards while outside verification is still pending. Two rival capability claims -- Zhipu's GLM-5.3 cybersecurity benchmark and OpenAI's own tripled ARC-AGI-3 score -- both ran into the same problem: the number that shipped first didn't survive a second look. OpenAI also launched ChatGPT for Teens the same week it expanded ads into Europe, and Cerebras unveiled a new inference chip claiming a 30x speed edge over GPUs.

What is The Harness ?

A daily summary of what is interesting and happening in the AI industry, with a focus on what this means for people building harness experiences that are used.

Good morning, it's Wednesday, August nineteenth.

In today's briefing we see OpenAI pausing its largest frontier training run over an unconfirmed cybersecurity threshold, two rival capability claims from Zhipu and OpenAI itself that didn't survive a second look, and new research suggesting reinforcement learning's reasoning gains come from reranking what models already know rather than teaching them something new.

First up - Today in the big model news;

OpenAI
OpenAI has paused reinforcement learning training on its next deployment-bound models for two weeks, covering tool use, code execution, and network access work specifically. The pause followed preliminary signs that its Astra model may clear the company's own Critical tier on its cybersecurity capability scale, meaning it could independently chain novel exploits against hardened systems. The lab's single largest planned frontier training run stays on hold past the two weeks, pending internal restart conditions. New safeguards include stricter sandboxing, cut standing privileges, and token level monitoring covering roughly a fifth of inference compute. And here's the catch: this is OpenAI grading its own homework. METR and Redwood Research, the outside evaluators named to verify results like this, haven't published anything yet. Treat vendor capability timeline claims as provisional until an outside evaluator weighs in, not as delivery dates.

That same problem showed up twice more. Zhipu's GLM-5.3 claimed a narrow win over Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol on a cybersecurity benchmark called CyberGym, but it was a single ungraded run with a rescaling factor applied inconsistently, and on the harder version of the test, the one requiring a working exploit rather than a spotted bug, GLM-5.3 trails both rivals by roughly twenty four points. Separately, two non-default API settings roughly tripled GPT-5.6 Sol's score on ARC-AGI-3, pushing it past Claude Opus 5's official result; ARC Prize's Francois Chollet says that's only fair if the settings are disclosed and available to everyone. Both claims follow the same pattern as OpenAI's own cybersecurity pause: the vendor's number ships first, unrefereed. Read the harder benchmark, not the press release headline.

OpenAI also launched ChatGPT for Teens, automatically routing younger users into a version with content restrictions, a Study Mode that guides rather than answers, and parental controls, following lawsuits alleging chatbots contributed to teen suicides, years after teens had already been using the adult product. The tier reads less like a safety feature than a distribution defense: parents and schools are the real gatekeepers, and the goal is keeping them from blocking ChatGPT or steering kids toward a rival that looks safer. Study Mode is the one piece built around an actual habit: assignment, guided walkthrough, grade. In the same stretch, OpenAI expanded ChatGPT Ads to thirty one European markets and gave investors downbeat profitability guidance, running a trust building product alongside a revenue product funding the company. Whether Study Mode's habit locks in before a rival looks safer decides whether this launch pays for itself.

In local model developments, a paper making the rounds, Akgul's ReasonMaxxer, argues that reinforcement learning post-training changes only one to three percent of a model's output tokens, concentrated at high entropy decision points, with the promoted token almost always already sitting in the base model's top five alternatives. The authors built an RL free method that matches full RL performance at roughly one thousand times lower compute. Commenters called the always-top-five claim statistically implausible for genuinely uncertain distributions, but the weaker version survives: reinforcement learning looks less like teaching a model to reason and more like reranking what it could already produce. A separate result points the same direction: BDH-CQ, a hundred and fifty million parameter model using latent space reasoning and temporary memory, hit twenty nine point five percent pass at two on ARC-AGI-1 for roughly seven hundredths of a cent per task, capability that size wasn't supposed to reach without a scratchpad trick doing most of the work. If reinforcement learning gains really are this narrow, a model that tests smarter on one benchmark may just be better tuned for that rerank, not more capable in general.

In the harness, tools and orchestration world, the same reframing, measure the harness rather than just the model, showed up twice more. Omar Sarvari's analysis of agent skills found they help mostly through procedural anchoring, about sixty six percent of the measured benefit, rather than supplying new facts, under five percent: a skill mostly tells a model how to do a known task, not what to know. A GitHub search turned up roughly three point eight million SKILL.md files already in the wild, evidence the format is becoming a packaging standard before any lab has tried to own it. Evaluation tooling moved the same direction: Hamel Husain's eval skills plugin converts raw model outputs into annotated failure modes, and Agent Arena is now filtering its rankings using signal from one point seven million real agent sessions instead of one-shot prompts. None of it needed a bigger model to answer what actually gets work done.

In AI Infra, Cerebras launched CS-4, a rack scale inference system it says delivers up to thirty times more tokens per second per user than GPU based systems, and ten times better throughput per watt than its prior generation. It's sampling with a small group of customers now. Cerebras isn't chasing Nvidia on training scale. It's betting that inference latency at the individual user level, not raw training compute, is the constraint that matters. That's a different bet than the compute landlord consolidation running through the rest of this year's infrastructure news: Nvidia financing OpenAI's data center campus, Groq folding into Nvidia's neocloud. Cerebras is the one vendor still betting an architecture bypass beats renting more GPUs. Whether a named enterprise customer, not just an unnamed small group, discloses a real production workload is what would confirm the bet.

That's the briefing. Have a great day, and don't forget to subscribe.