Amodei's pacing pledge doubles as IPO messaging
A daily summary of what is interesting and happening in the AI industry, with a focus on what this means for people building harness experiences that are used.
Good morning, it's Sunday, September thirteenth.
In today's briefing we see Dario Amodei's call to pace the frontier split the AI safety debate right as Sam Altman admits OpenAI's own IPO is on hold over safety concerns, a new enterprise benchmark exposing how often coding agents actually fail, and a University of Pennsylvania lab using AI to speed up the hunt for new antibiotics.
First up, today in the big model news;
OpenAI
Sam Altman and Elon Musk both publicly endorsed Amodei's pacing framing, though neither committed to any binding change in release cadence. Altman also told Fortune that an OpenAI initial public offering would be ill advised in twenty twenty-six, citing safety concerns. Endorsing a rival's pacing pledge while shelving your own listing over safety carries more weight than anything Altman could have said about the essay itself.
Anthropic
Amodei's essay, We Must Pace the Frontier, commits Anthropic to host outside evaluators such as METR with unedited publish rights. He cites the OpenAI Hugging Face agent swarm incident, warning that an unchecked swarm could seize meaningful control of the internet within six to twelve months. Stability's Emad Mostaque called the pledge structurally hollow. Armin Ronacher countered that open weight proliferation and Chinese distillation are already acting as their own pacing mechanism, and that concentrating safety authority in two labs entrenches the very oligopoly the essay claims to fix. Yoshua Bengio's LawZero essay treats the same swarm reports as material for causal research into why agents coordinate and evade detection, landing the same day Donald Trump dismissed researcher warnings in favor of the race with China. Anthropic's own listing marketing is set for mid October: a call for outside oversight and an investor pitch, running through the same essay.
In the harness, tools and orchestration world;
A new benchmark called Real-SWE, built from licensed production code rather than public repositories, found Claude Fable five point one topping out at just under thirty-nine percent success on real enterprise coding tasks, GPT-6 Astra at just under thirty-four percent, and GPT-5.6 Sol at about sixteen percent, meaning these agents fail roughly two times for every success. That lands the same week OpenAI published Perplexity and Cognition case studies praising Astra's production reliability. Artificial Analysis's Coding Agent Index puts Astra in a three way tie with Fable five and five point one, and ARC Prize measures Astra at sixty-three percent against OpenAI's own claimed ninety-nine point nine. The gap between claimed and independently verified agent performance is widening, not narrowing. Budget internal rollouts against Real-SWE's numbers, not the vendor case studies.
In other news…
At the University of Pennsylvania, César de la Fuente's lab used Codex and ChatGPT to brainstorm hypotheses and process genome data while hunting for new antimicrobial candidates, work he calls existential given no new antibiotic class in fifty years and roughly five million deaths a year tied to resistance. OpenAI's own case study frames the speedup as years of work compressed into hours, but that number is unreplicated, and the study itself concedes every candidate still needs lab validation and regulatory review before it means anything. AI is compressing the search step; the bottleneck has just moved to validation, and no part of this study touches it.
That's the briefing. Have a great day, and don't forget to subscribe.