The Harness

AMD buys its way into inference silicon

Show Notes

AMD is acquiring Taalas, a startup that etches model weights directly into silicon, the clearest sign yet that custom inference chips are consolidating through M&A rather than staying a venture bet. A public game found humans miss one in three threats when approving AI agent commands, undercutting human review as a real safety net just as Meta pushes contested self-reported benchmark claims for Muse Spark 1.2 and OpenAI ships an open agent-plugin standard five major vendors agreed to jointly. Anthropic also quietly cut Fable 5's biology-safety false positives by 85 percent, showing how launch-day caution gets walked back once real usage data comes in.

What is The Harness ?

A daily summary of what is interesting and happening in the AI industry, with a focus on what this means for people building harness experiences that are used.

Good morning, it's Friday, August seventh.

In today's briefing we see AMD buying its way into custom inference silicon, a public safety game showing humans miss a third of malicious agent commands, and OpenAI shipping an open plugin standard for agent skills that five major vendors signed onto together.

First up - Today in the big model news;

OpenAI
OpenAI folded its GPT-5.6 Sol and Luna models into a single family with a reasoning-effort slider, and gave free-tier users unlimited access to Luna chats. That's one adjustable model line taking the place of separate branded tiers.

Anthropic
Anthropic rewrote the safety classifier behind Fable 5's biology guardrails, cutting false-positive reroutes to the weaker Opus 4.8 by about eighty five percent, so lab-result interpretation, symptom research, and routine clinical questions now get answered directly instead of getting blocked. It's the predictable second half of a launch pattern: ship an overly cautious classifier on day one to avoid a bio-risk headline, then loosen it once real usage data shows what got over-blocked. Expect the same recalibration on Fable 5's other gated domains, cybersecurity, chemistry, and distillation queries, as false positives show up in support tickets rather than safety reviews.

Meta
Meta's Muse Spark 1.2 claimed gold-medal-equivalent scores across five STEM Olympiads in physics, chemistry, and math with no external tools, a top five finish on the Vals Index at sixty nine cents a test, and the first result over sixty percent on Finance Agent version two. Outside analysts are skeptical: the roughly eighty three percent Terminal-Bench score Meta is touting comes from Meta's own harness with no independently verified leaderboard entry, Meta's prior model missed its own claimed score by nearly four points once outsiders checked, and two of three rival comparisons used the competitor's second tier model rather than its flagship. A model competing on price is now also competing on self-graded benchmarks, and both claims get equal billing in the coverage regardless of which one holds up.

In the harness, tools and orchestration world;

OpenAI shipped Agent Plugins 1.0, an open, vendor-neutral packaging format for agent skills and MCP server configs, built jointly with AWS, Cursor, GitHub, Microsoft, and Vercel, governed by a public technical steering committee so no single company controls the roadmap. ChatGPT, Codex, Cursor, Copilot, Kiro, and VS Code all support it at launch. Five competitors agreeing to one plugin format in a single move is a bigger commitment than any one lab's model release: it says the packaging layer for agent skills is now infrastructure worth standardizing jointly, the same instinct behind Nvidia's NOOA and OpenAI's own Codex Security push in July. It's also a bet that durable value lives in the harness rather than the model underneath, since the model still has to be rented either way.

A public game testing whether people can catch compromised AI agent commands, ordinary calls like git status mixed with malicious ones like reading AWS credentials, found players missed one in three threats across forty thousand runs and four hundred nine thousand decisions, with a third of sessions net negative once missed threats and false blocks were tallied. It's the third counterweight in a week to the harness-loop thesis, after Cloudflare's provenance-audit layer and RufRoot's default security gaps: the accumulated-trust layer that makes a harness sticky is also where it fails, this time in the human fallback teams lean on when automated governance isn't built in. Treat any product that ships human-approval gates as its stated safety mechanism as unverified until it publishes a real accuracy number instead of a marketing claim.

In AI Infra

AMD is acquiring Taalas, a Toronto startup that etches model weights directly into silicon instead of storing them in HBM memory. Its first test chip served Llama 3.1 8B at roughly seventeen thousand tokens per second on a six nanometer process, terms are undisclosed, and the deal closes in the fourth quarter, with Taalas folding into AMD's Helios racks alongside Instinct GPUs and EPYC CPUs. Etched raised three hundred million dollars in July on the same bet, that inference economics reward hardware frozen around one model, except this time a chip incumbent bought the approach outright instead of a startup raising around it. That's custom inference silicon moving from venture bets toward acquisition-driven consolidation. The tradeoff is real: an etched chip can't be repointed at a new model release, so it only pays off for workloads stable enough to freeze in place.

Quick hits from the consumer side;

Google Maps now lets its AI agent book hotels and order food directly in the app instead of routing through a third party. Suno is adding watermarks to AI generated songs as its copyright litigation escalates. OpenAI's first consumer hardware product, a smart speaker, is reportedly landing at three hundred to four hundred dollars.

That's the briefing. Have a great day.

Hi, this is Jamie. Thanks so much for listening to The Harness. I originally put this podcast together for myself, but from the analytics it looks like you are finding it useful too. If you haven't already, please go ahead and hit subscribe or save. It really helps me and the show.