The Harness

Switzerland enters the sovereign AI stack

Show Notes

Switzerland's EPFL and ETH Zurich released Apertus, the first foundation model stack with full EU AI Act compliance documentation, directly addressing enterprise demand for sovereign alternatives after the Fable-Mythos ban. Sakana AI's Fugu multi-agent API claims to match Opus 4.8 and GPT-5.5 by dynamically routing to specialist agents, raising the question of whether orchestration layers are now the real procurement decision. A critical OpenAI Codex logging bug is writing up to 37TB to developer SSDs in 21 days — a silent hardware risk not mentioned in any onboarding material.

What is The Harness ?

A daily summary of what is interesting and happening in the AI industry, with a focus on what this means for people building harness experiences that are used.

Good morning, it's Monday, June twenty-second.

In today's briefing we see Switzerland's EPFL and ETH Zurich clearing the EU AI Act compliance bar with Apertus, Sakana AI's multi-agent orchestration layer matching frontier benchmarks, and an issue with OpenAI Codex creating hardware risks.

First up - Today in the big model news;

Google and Deepmind / Gemini

DeepMind released ATLAS, a system that generates interpretable mechanistic models from experimental data and autonomously designs follow-up experiments to test them. This differs from pattern-matching approaches like AlphaFold or the Co-Scientist multi-agent work: ATLAS generates hypotheses and runs experiments to falsify them. It shifts AI-for-science from "find patterns in data" to "generate and test theories," reframing how research institutions integrate AI into wet-lab workflows. The hire of Anthropic's John Jumper makes more sense in this context; Jumper's credibility in scientific reasoning provides the foundation for advancing this kind of hypothesis-driven automation. For research teams adopting agentic AI into experimental science, the architecture choice now includes hypothesis generation as a first-class capability, because ATLAS treats theory building as an autonomous process that can drive follow-up experiment selection.

In local model developments;

Switzerland's EPFL, ETH Zurich, and CSCS released Apertus 8B and 70B this week, the first foundation models explicitly documented to satisfy EU AI Act requirements: PII removal, opt-out functionality, memorization prevention, and complete training data transparency. Swisscom adopted Apertus as the base for its Swiss AI Platform. These aren't frontier-class models, but they're auditable and competitive at scale, with a complete paper trail from data to weights to alignment principles that enterprises can hand to a data protection authority. For enterprises in regulated industries, auditable sovereign alternatives are now available within the EU governance perimeter, because the Fable and Mythos ban demonstrated that dependence on US-based frontier access creates both legal and operational exposure.

Open-weight adoption has crossed into pragmatic procurement territory. OpenRouter data shows sixty percent of traffic now on open weights, up from forty percent in March, a shift framed not as advocacy but as risk management. The toolchain is mature enough that migration costs have fallen. For AI PMs still evaluating closed versus open, stress-test your fallback posture now, because the Fable and Mythos restrictions demonstrated that frontier intelligence access can be revoked abruptly and the next supply shock will arrive without warning.

In the harness, tools and orchestration world;

Sakana AI released Fugu and Fugu Ultra, an API-delivered multi-agent orchestration system backed by two ICLR 2026 papers on evolved coordinators and natural-language coordination strategies. Fugu Ultra claims to match or exceed Opus 4.8, Gemini 3.1 Pro, and GPT-5.5 on coding, reasoning, and scientific benchmarks via a single OpenAI-compatible endpoint, eliminating SDK migration. For product teams evaluating procurement, expect the decision to shift from model selection to orchestration layer selection, because multi-agent coordination via orchestration is achieving frontier-level performance and the vendor lock-in risk moves up the stack.

A memory framework called AtomMem treats persistent facts as first-class objects for long-lived agents, addressing the KV-cache degradation and session compression challenges that emerge when agents run for hours or days. This pairs with emerging concerns that agentic benchmarks remain unreproducible: scores vary across runs, tool permissions differ between reported and actual conditions, and timeout policies go unstandardized. When a model claims a performance level on a coding benchmark, the real-world equivalent depends entirely on harness choices. For teams migrating from frontier APIs to open models, validate performance on your own harnesses before committing, because benchmark gaps discovered after deployment create unexpected production risk.

An operational warning: OpenAI Codex's default logging writes approximately thirty-seven terabytes to local SSDs in twenty-one days, roughly six hundred and forty terabytes per year, exceeding consumer drive write endurance in under a year. The mechanism is write amplification from continuous logging of protocol noise. This isn't surfaced in any onboarding material. For developers shipping Codex workflows, check SSD health on machines running Codex as a daily driver and adjust the logging configuration, because the default setting creates silent hardware degradation.

That's the briefing. Have a great day.