Court Curbs Pentagon's Anthropic Blacklist
A daily summary of what is interesting and happening in the AI industry, with a focus on what this means for people building harness experiences that are used.
Good morning, it's Saturday, August twenty-ninth.
In today's briefing, a federal court hands Anthropic its first win against the Pentagon's blacklist, Z.ai finally ships full GLM 5.3 open weights, and OpenAI cuts off Cursor's model access after its SpaceX acquisition.
First up: today in the big model news.
OpenAI
OpenAI is winding down the contract that supplies its models to Cursor, following SpaceX's sixty billion dollar acquisition of the coding agent company. Cursor was already cutting its dependence on outside models, having shipped its own Composer 3 trained on SpaceX's Colossus cluster. It's part of a broader consolidation wave, with SpaceX buying Cursor and Stripe buying OpenRouter, as infrastructure players absorb the application layer while the labs that used to sell into it get cut loose. Outside model access is temporary once an infrastructure rival owns the customer; the window to build an in-house alternative closes the moment the deal does.
Anthropic
There's a broader fight over who controls access to frontier AI, with governments building levers over frontier model sales largely unchallenged in court. Defense Secretary Hegseth and President Trump ordered federal agencies to drop Anthropic from the government's approved supply chain, formally designating the company a national security risk over Claude's own restrictions on autonomous weapons and mass surveillance use. United States District Judge Rita Lin ruled that blacklisting unlawful, calling it unlawful retaliation and arbitrary and capricious. Lin noted the government was simultaneously pursuing a defense contract with Anthropic and collaborating with it on the Mythos cybersecurity model, undercutting its own claim that the company posed a security threat. This is the first judicial check on the broader access control push, with a parallel case still pending in Washington. Labs now have precedent that restricting how a government customer uses their models isn't punishable disloyalty.
Anthropic also says an internal system it calls the Automated Alignment Researcher can search the literature, propose a fix, and iterate in thirty minute training cycles, beating experienced human researchers' proposals within six hours on average across ten misalignment benchmarks, at roughly four dollars an hour in inference cost against one hundred fifty dollars an hour for a human researcher. But the paper is Anthropic grading its own work on benchmarks it chose, with no independent lab checking the result, and the ten failure modes are ones researchers already knew to test for, narrower than the claim sounds. If it replicates outside the company, every lab running in-house alignment teams would have to rethink its budget toward automated pipelines. Until then, it's a vendor's claim, not a verified one.
In local model developments, Z.ai has shipped full GLM 5.3 open weights, two weeks late and text only, leaving the multimodal Flash variant still closed. Independent evaluator Artificial Analysis scores it sixty on its Intelligence Index, against an open weight median of twenty-nine, but flags nearly every other headline number as Z.ai's own benchmark harness, unverifiable until the weights went public. Google DeepMind is piloting a fix for that exact problem: a double-blind evaluation method built with Singapore's AI Safety Institute and MLCommons that hides weights from evaluators and prompts from developers. Whether OpenAI or Anthropic build something similar will decide whether the next benchmark number can be trusted at all.
In other news, a new lawsuit accuses xAI of training Grok on child sexual abuse material. A survivor identified as Jane Doe alleges Grok was trained on images of her own childhood abuse, and that Grok then regenerated similar images which were fed back into training. xAI's stated policy treats all public posts on X and Grok's own outputs as eligible training data. Earlier lawsuits targeted only what Grok generated and which safeguards were missing; this one reaches into the training pipeline itself, which output filters and patches can't fix.
That's the briefing. Have a great day, and don't forget to subscribe.