Coding agents' valuations outrun their own multiples
A daily summary of what is interesting and happening in the AI industry, with a focus on what this means for people building harness experiences that are used.
Good morning, it's Thursday, August thirteenth.
In today's briefing four frontier models, Grok 4.6, DeepSeek V4-Pro, Qwen 3.8 Max, and Microsoft's MAI-Thinking-1, all shipped within days of each other, Google folded Gemini directly into its new Pixel hardware lineup, and a security researcher found attackers spoofing AI crawler traffic to slip past site defenses.
First up - Today in the big model news;
Alibaba
Qwen 3.8 Max, a two point four trillion parameter open model, finally got the infrastructure treatment that credibility requires: day-zero support in vLLM, and Unsloth's dynamic quantization work shrank its four point nine terabyte checkpoint down to three hundred ninety-seven gigabytes. That compression matters more than any benchmark table, because it decides who can actually run a model this size outside a hyperscaler's data center.
Google
Google's Made by Google event bundled Gemini into the Pixel eleven lineup, the Pixel Watch five, and a new Pixel Tag, the same week Gemini crossed one billion monthly active users, with sixty-three percent of interactions arriving by voice through Assistant, Android, and Search defaults. That user count is a distribution result more than an adoption one: bundling Gemini into three device categories just extends the same default placement past the phone. Owning the operating system, the assistant, and the device at once is harder for OpenAI or Perplexity to contest than a better model alone.
In other lab news today, three more frontier models shipped within days of Qwen's update, all racing to strip pricing power away from whoever ships the most expensive model. xAI's Grok 4.6 launched with a five hundred thousand token context window and a claim of matching GPT-5.6 Sol on the Artificial Analysis Intelligence Index, at the same price as its predecessor: a capability upgrade sold as a coding and agent tool, not a price cut. DeepSeek pushed V4-Pro fully out of preview the same week, a one point six trillion parameter mixture of experts model with agent benchmarks approaching several Western rivals, priced unchanged even as DeepSeek keeps warning developers that a significant hike is coming after its cheap tier got overwhelmed by demand. Microsoft's MAI-Thinking-1 rounds out the cluster more quietly, its first in-house reasoning model built around tool use rather than benchmark maximization. DeepSeek is holding its price because compute, not strategy, is the constraint; xAI is holding its price because it already claims parity without needing to cut.
In local model developments, LLM Compressor's new REAP expert pruning technique and vLLM's Azure Blob support for model loading are both aimed at making these huge sparse models cheaper to serve, not just cheaper to train. Lightricks' LTX-2.5 pairs video generation with synchronized audio in one pass, and Cohere's North Micro Vision targets document understanding specifically rather than general vision: both are bets that the next competitive axis is task specific quality, not another general benchmark chase. The frontier labs are converging on raw capability; the fight over who becomes infrastructure is happening one layer down, in serving cost and task fit.
In AI Infra, researcher Gavin King found a surge in vulnerability scanning across thousands of sites, attackers spoofing AI crawler user agents like ClaudeBot and Googlebot to slip past defenses built to let AI crawlers through, with bursts up to seventy thousand requests a minute from Google Cloud IP space targeting exposed credential files. It's a familiar failure: an assumed trust abstraction, here the web's crawler allowlisting convention, becomes the cheapest attack surface once nobody actually verifies it. Checking reverse DNS and IP ranges against the real crawler is cheap and already exists, and it stays cheaper to exploit than to fix until a scan like this costs someone real money.
In other news, Cognition is in talks to raise at a forty billion dollar valuation, up from twenty-six billion three months ago, with revenue nearing one billion dollars, double its last raise figure. That growth actually compresses the multiple investors are paying, from about fifty-two times revenue down to roughly forty times. Blacksmith, a code testing startup, jumped to five hundred fifty million dollars from sixty million in under a year, and Lovable confirmed a thirteen point three billion dollar valuation on a fresh four hundred million dollar round. Investors are pricing continued growth here, not re-rating on hype, which raises the bar for coding tool startups still pitching on demo quality.
Separately, Amazon updated Twitch's settings so streams, clips, chat, and channel text and images train its generative AI models by default, requiring streamers to opt out; a support forum comment demanding opt-in drew nearly fourteen thousand upvotes within hours. Twitch is a live content corpus no rival can scrape off the open web, so an opt-out default captures it before streamers organize. This is a competitive move, not a regulatory one: whichever rival advertises opt-in only first turns creator trust into a recruiting pitch, an opening Amazon's default just handed away.
Elsewhere, a security test found AI agents leaking sensitive data more often than they blocked injected prompts, and the debate over whether Claude's new invisible watermarks survive paraphrasing continues unresolved, now showing up as real user pushback rather than just a technical argument, since people worry the watermark could expose their AI use at work. Both point to the same gap: safety features are shipping faster than anyone is verifying they hold up under normal use.
That's the briefing. Have a great day, and don't forget to subscribe.