The Harness

Nemotron 3.5 undercuts the frontier-agent tier — until you check the compare

Show Notes

Researchers found that OpenAI, Anthropic, and Google's encrypted chain-of-thought blocks are interchangeable across sessions and models, letting a weaker jailbroken model decode a stronger one's hidden reasoning in plaintext and leaking thousands of API keys and passwords in the process. NVIDIA's new open Nemotron 3.5 Lightning model shipped with a viral claim that a legal-AI vendor's fine-tune beat Claude Opus 4.6 on their own benchmark, a claim that falls apart under a check of the vendor's own prior write-ups. OpenAI's only dedicated ethicist left without being replaced, joining a summer of safety-leadership departures right as the company asks regulators to trust its own self-graded call on a frontier model's cyber capability.

What is The Harness ?

A daily summary of what is interesting and happening in the AI industry, with a focus on what this means for people building harness experiences that are used.

Good morning, it's Wednesday, August twelfth.

In today's briefing we see researchers cracking open the encrypted reasoning traces that OpenAI, Anthropic, and Google all rely on, NVIDIA's new open Nemotron 3.5 Lightning model catching a legal AI benchmark claim that doesn't survive a check, and another senior safety departure at OpenAI right as it asks regulators to trust its own grading.

First up - Today in the big model news;

OpenAI
OpenAI's only dedicated ethicist, Chloé Bakalar, left the company in July after less than a year, and the role hasn't been refilled. OpenAI says ethics is now handled across teams, with no single position responsible for it. She's the fourth senior departure in about a month, following safety leads Johannes Heidecke and Joshua Achiam and eight year veteran Brad Lightcap, who announced this week he's leaving to build something new. The exits land as OpenAI asks regulators to trust its own self graded call that its Astra model can't be ruled out for critical cyber capability. The teams built to check a lab's self grading are thinning at the exact moment that grading carries the most weight.

Google
Google's Gemini app crossed one billion monthly active users this week, the fastest growing product in the company's history by its own count, matching the milestone ChatGPT hit back in June. The two got there differently: sixty three percent of Gemini's interactions run through voice, meaning Google converted placement it already owned, in Assistant, Android, and Search, into Gemini usage, while ChatGPT built its billion from a standalone app with no operating system default to lean on. That gap shows up in how each funds its free tier: OpenAI expanded ChatGPT ads, already live in the United States since July, to the UK, Mexico, Brazil, Japan, and South Korea this week, while Gemini's free tier stays ad-free, backed by the ad business Google already had. Clearing a billion users stops being a moat the moment two labs both do it; the next contest is which one needs its own users to pay for the free tier.

Anthropic
Anthropic shipped imperceptible statistical watermarks into Claude's text output today, plus signed provenance metadata on generated files: a keyed sampling bias during generation that a detector can recover later with a statistical test. It's a reasonable compliance move, but the community's first reaction was the right one: paraphrase the output through a second model and the statistical signal likely doesn't survive. Provenance is becoming a checkbox every lab has to ship, without any of them yet making it stick.

Local model developments
NVIDIA released Nemotron 3.5 Lightning, an open licensed mixture of experts model with thirty one point six billion parameters and three point six billion active, and inference vendors including Together AI, Ollama, Baseten, vLLM, and Perplexity had it running within a day. NVIDIA never collects model license revenue, so commoditizing the model layer keeps GPU burning workloads growing instead of consolidating behind a few closed APIs. Harvey AI's post training exercise looked like the sharpest proof of that trade: its Legal Agent Bench score moved from zero to about eight percent, reportedly beating Claude Opus 4.6 at a third of the output tokens. That claim doesn't hold up: a comparable write-up on the related Nemotron 3 Ultra model puts the real score around six percent, landing between Anthropic's Sonnet 4.6 and Opus 4.6, and Opus's own cited score moved from about four percent to about six and a half percent between Harvey's two reports. Treat the throughput gains as real and the specific claim of beating Opus as unconfirmed until Harvey publishes a model specific result.

In AI Infra
Researchers at the ELLIS Institute Tübingen and the Max Planck Institute found that the encrypted reasoning traces OpenAI, Anthropic, and Google return to clients, so they can be replayed in follow up requests, are interchangeable across sessions, users, and even models within the same provider. Feed a stronger model's encrypted trace into a weaker, jailbroken model on the same account, and it decodes and repeats the reasoning in plain text. The recovery run pulled roughly seven thousand public traces, several dozen leaked API keys, and a scattering of emails and passwords people had pasted into prompts, a real security problem even before the more contested claim that Kimi may have been trained or distilled on reasoning extracted this way. One analyst pushed back that the encryption was built as a stateless distributed inference optimization, not a confidentiality guarantee, so calling this practical mass theft of training data overstates the demo. Both things can be true: the traces were never built to survive a determined adversary, and providers now have to decide whether to encrypt for real or stop implying the reasoning is hidden at all.

On the regulatory front today, Senator Bernie Sanders sent a formal letter to Altman, Amodei, and Zuckerberg urging an immediate pause on frontier development, naming loss of control risk, bioweapon enablement, and model escape scenarios directly. It lands one day after House Democrats asked the same three CEOs to testify under oath about AI driven hacks. Neither move has produced anything with teeth yet, but a formal pause letter and a same week hearing request aimed at the same three companies is the most coordinated congressional pressure the labs have faced together this year.

Quick hits from the consumer side;
OpenAI shipped a ChatGPT desktop app for Linux, covering Ubuntu, Debian, and Fedora on both x64 and ARM64, with project import from other agent platforms built in.
Spotify said it will start labeling AI generated artist profiles next month, with a detection plan for creators who skip self-disclosure.
xAI put Grok Bot into beta, an AI teammate that runs from its own persistent cloud computer, signs into a user's existing tools, and keeps memory across tasks.

That's the briefing. Have a great day.

Hi, this is Jamie. Thanks so much for listening to The Harness. I originally put this podcast together for myself, but from the analytics it looks like you are finding it useful too. If you haven't already, please go ahead and hit subscribe or save. It really helps me and the show.