The Harness

A Fields Medalist bets his career on AI accelerating math 100x.

Show Notes

A Fields Medalist takes leave from academia to work on AI safety at OpenAI, one day after Astra's machine-verified math proofs, while Alibaba's Qwen3.8-Max finally gets outside benchmark numbers two weeks after shipping with none. Google quietly kills its AI Studio mobile app in favor of building apps straight out of Gemini conversations, and a maximum-severity vulnerability in the open-source Ruflo agent platform shows the same harness memory that makes agents sticky can also be hijacked with one unauthenticated request. Plus: an arxiv paper finds AI-migrated COBOL passes its checks while still shipping new bugs.

What is The Harness ?

A daily summary of what is interesting and happening in the AI industry, with a focus on what this means for people building harness experiences that are used.

Good morning, it's Monday, August third.

In today's briefing a Fields Medal winner is leaving academia to work on AI safety inside OpenAI, Alibaba's Qwen3.8-Max finally has independent benchmark numbers behind it, and a maximum severity vulnerability turned up in a popular open source agent platform.

First up - Today in the big model news;

OpenAI
OpenAI revealed Astra, a research effort behind ten machine-verified Lean four proofs of decades-old open math problems. One day later, Jacob Tsimerman, this year's Fields Medal winner for his work on the André-Oort conjecture, announced he's taking leave from the University of Toronto to work on AI safety inside OpenAI. He's betting that AI will soon accelerate math research a hundredfold and could pose a real safety risk once it does. It's a notable pairing: a credentialed outside mathematician choosing to verify frontier math capability from inside the lab making the claim, rather than fact check it from outside the way Terence Tao did back in July. Whether that counts as real verification or just a well credentialed hire depends entirely on what he's allowed to publish.

Alibaba
Alibaba's Qwen3.8-Max finally has real benchmark numbers behind it, two weeks after the company previewed the two point four trillion parameter model with nothing but a social media post calling it second only to Fable five. Independent testers now put it at eighty six point six on Terminal-Bench two point one, ahead of GPT-5.6 Sol and Fable five on several coding evals. Alibaba still hasn't published an official benchmark table or model card, so this number exists purely because outside testers went and got it themselves. When the strongest evidence for a flagship model comes from testers the vendor didn't authorize, disclosure becomes the thing worth watching as closely as the score.

Google
Google pulled its standalone AI Studio app from both app stores the day before launch, despite more than eight hundred thousand pre-registrations, and is folding its app generation features directly into the main Gemini app. Google is betting that convenience wins over discoverability here: reaching everyone who already has Gemini open matters more to them than the store listing and engagement data a standalone app would have produced.

In AI Infra
Noma Labs disclosed a maximum severity vulnerability, which researchers dubbed RufRoot, in Ruflo, a popular open source agent orchestration platform. Its MCP Bridge exposed two hundred thirty three tool endpoints over HTTP with no authentication, letting a single request achieve full remote code execution, steal the agent's API keys, and rewrite its stored memory to influence future sessions. Ruflo patched within twenty four hours. The same accumulated memory that makes a harness sticky is, by default, a single unauthenticated point of failure for every team running it. If you're deploying an agent harness, treat who can reach your MCP bridge as a standing question, not a one time setup checkbox.

In other news...
Andrej Karpathy's most discussed post of the day argued that informal one shot benchmarks are aging out. He gave Claude Opus five the opening paragraph of Lord of the Rings, a million token budget worth about ten dollars, and asked for a procedural three dimensional render; two hours and fifty five hundred lines of JavaScript later it produced a janky but genuinely functioning rendered world. His point is that tests like drawing a pelican on a bicycle measured a capability tier models have already cleared, and nobody has built the equivalent gut check yet for open ended, long horizon generation. That's a flag that benchmark design hasn't kept pace with what these models can now attempt.

A new arxiv paper found that AI models migrating legacy COBOL code to Java produced output that passed every migration check while still introducing behavioral bugs the checks never tested for. The code worked by the metric it was graded against, and the real defects surface later, in production behavior nobody scoped the checklist to catch. If you're budgeting a COBOL migration as a guaranteed AI win, build a behavioral equivalence pass into the plan from the start.

That's the briefing. Have a great day.

Hi, this is Jamie. Thanks so much for listening to The Harness. I originally put this podcast together for myself, but from the analytics it looks like you are finding it useful too. If you haven't already, please go ahead and hit subscribe or save. It really helps me and the show.