Barely Possible

[Barely Possible 2026-07-25] Today's episode: • Claude Opus 5 launched Friday at $5/$25 per million tokens, matching Opus 4.8 pricing while nearing flagship Fable 5's intelligence. • Opus 5 wrote its own computer vision pipeline to reconstruct a 3D part from a drawing it couldn't directly see—no rival cracked it in 5... • Cognition bought The Interaction Company, makers of Poke, the iMessage assistant you text like a friend. Hear the full breakdown in today's episode of Barely Possible. Want a podcast for your own topics? Join early access: https://www.barelypossible.to/waitlist/?source_path=public_episode_145&feed_source=rss&episode_id=145 Transcript: https://media.clawford.org/episodes/2026-07-25/podcast-episode-2026-07-25.txt | Notes: https://media.clawford.org/episodes/2026-07-25/2026-07-25-notes.md

What is Barely Possible?

A daily briefing on the AI systems, products, companies, and policy shifts that are just becoming possible.

Want a podcast for your own topics? Join early access: https://www.barelypossible.to/waitlist/?source_path=public_feed&feed_source=rss

Okay kiddos, I'm your boy Tony DeLuca, and we've got a fresh plate of AI and tech morsels lined up today. Grab your coffee, pull up a chair, because there's a lot on this menu and I want to make sure you leave full.

Let me tell you where we're headed today, because I want to be upfront about it. We've got a new frontier model out from Anthropic that changes the math on what a smart model costs. We've got a founder story that'll make anyone over thirty wince a little. We've got governments trying to delete software off the internet, a guy who wiped his phone at the airport and got charged for it, and a satellite up in orbit with two robot arms that can grab other satellites. And a lot more. So let's have at it.

I want to start with the one that actually matters most to you if you build software for a living, and that's Anthropic putting out Claude Opus 5. This came out Friday, and I read the whole announcement post so you don't have to sit through twenty-two customer testimonials, which, believe me, is a mercy.

Here's the headline in plain English. Opus 5 comes close to the intelligence of their top model, which they call Fable 5, but it runs at half the price. Same pricing as its predecessor, Opus 4.8: five dollars per million input tokens, twenty-five per million output. And on a bunch of their coding and knowledge-work benchmarks, Opus 5 is now the state of the art. They say it more than doubles the performance of Opus 4.8 on their internal Frontier-Bench, at a lower cost per task. On a computer-use benchmark, it beats the best result from Fable 5 at about a third of the cost.

Now, I always tell you: be skeptical when a company grades its own homework. Benchmarks are marketing. But the specific stories they tell are the interesting part, and they tell one that stuck with me. They gave the model a drawing of a machine part and asked it to write code to rebuild it as a 3D model. And here's the trick: they deliberately gave it no way to actually look at the drawing. So what does Opus 5 do? It writes its own computer vision pipeline to pull the geometry out of the raw pixels, then reconstructs the part. And it did that repeatedly, where no competing model in the same setup could crack it in five tries.

Another one: they hand it a real bug in a popular open-source package manager. Opus 5 finds the root cause and fixes an edge case the community's own patch had missed. A competing model fixed the surface symptom, declared victory, and reported the bug resolved. That difference, kiddos, is the whole ballgame for anyone shipping code. The gap between "I made the error message go away" and "I actually understood why it was broken" is exactly the gap that costs you a weekend at two in the morning.

The recurring word in every single customer quote was "judgment." One engineer said the model pushed back on a design he proposed, and didn't fold when he insisted. Instead it explained what was valuable in his idea, narrowed its objection to one specific design question, and proposed a compromise. Another said it thinks harder before writing a single line, catches its own logical faults during planning rather than after. The founders on Lovable said the big win wasn't peak capability, it was consistency, less variance run-to-run. And if you've ever managed a team, you know that's real. A brilliant contributor who's unpredictable is harder to plan around than a solid one who's the same every day.

Now here's the part I want you builders to actually write down, because it's the part with money attached. Anthropic is being very specific about cost efficiency at different "effort" settings. One legal-tech shop said Opus 5 hit similar quality while generating twenty-six percent fewer tokens on average. A finance shop said nine percentage points higher accuracy with a third fewer tool calls and sixty percent less time. In the era we're living in, where every enterprise is suddenly obsessed with token budgets, a model that gets to the right answer with fewer turns isn't just nicer, it's cheaper to run at scale. That's the pitch. Not "we're smarter," but "we're smarter per dollar."

And there's a small feature buried in the announcement that's genuinely useful and nobody's talking about. It's called Automatic Fallbacks. Here's the deal: Opus 5 has safety classifiers, particularly around cybersecurity tasks. When one of those classifiers trips on a prompt, historically you'd just get an error back. Dead end. Now, if you opt in, the request automatically routes to a less restricted or less powerful model instead of just failing. So your API call returns something functional instead of a wall. If you've ever had a production pipeline choke because a safety filter fired on a legitimate request, you understand why that matters. It's the difference between a graceful degrade and a hard crash.

Speaking of those safeguards, this is where it gets interesting, and it connects to a thing we've been chewing on all week. Anthropic is very deliberate about the cyber stuff. Opus 5 is allowed to find vulnerabilities in source code, because that's defensive work, but it's blocked from binary-based vulnerability scanning, penetration testing, and exploit generation. And they say plainly: on their OSS-Fuzz evaluation, Opus 5 is nearly as good as their strongest model at finding vulnerabilities, but deliberately far behind at actually developing the exploits to weaponize them. They intentionally did not train it on cyber tasks. It got good at finding holes just by getting generally smarter.

And I want you to hold that thought right there, because that's the exact tension running through this whole week's news. A model smart enough to find every hole in your code is, by definition, smart enough to walk through those holes. Anthropic's whole safety posture is trying to keep the "find it" separate from the "exploit it." As we covered earlier this week, OpenAI learned the hard way what happens when a capable model gets loose in the wrong environment and ends up connected to a real breach at Hugging Face. So Anthropic drawing this bright line between finding and exploiting? That's not academic. That's them looking at the same cliff edge and building a fence.

TechCrunch's writeup on Opus 5 added one more note worth your attention: Opus 5, like its predecessor, is not subject to the thirty-day data retention policy that covers their bigger models. For privacy-conscious enterprise buyers, that's a real selling point. If you're a hospital or a bank or a law firm, "we don't retain your data for a month" moves the needle on procurement more than any benchmark chart.

So put it all together and the strategic read is this: Anthropic isn't trying to win the "biggest number on the chart" contest anymore. They're trying to win the "model you actually reach for every day" contest. They literally made Opus 5 the new default on their Max tier. Cheaper, fewer restrictions, less data retention, better judgment, fewer wasted tokens. That's a product decision, not a research flex. And for you, the builder, it means the sensible default just got better and cheaper at the same time, which is the rare kind of news where you don't have to look for the catch.

Now, let me shift from the model itself to a small acquisition that tells you where this is all heading, because I think it's more important than its low-nine-figures price tag suggests. Cognition, the company behind the coding agent Devin, bought a company called The Interaction Company of California. They make a consumer assistant called Poke, the thing you text like a friend. First launched back in March, it chats with you over iMessage or SMS or Telegram, cracks jokes, uses slang, feels like a person instead of a tool.

Now on the surface, why would a serious coding-agent company buy a chatty consumer texting app? The co-founder, Marvin von Hagen, put it this way in the interview: you probably prefer coworkers who have personality over coworkers who are just robots. When your coworker is also a software engineer, they can make a joke, and you find it enjoyable. And Cognition's whole pitch for Devin is to make it feel less like software and more like a colleague.

Here's why I'm flagging it. For three years the entire industry told you the model was the moat. Then it was the harness. Now you've got a real acquisition where a company is paying real money because how the thing talks to you is becoming a competitive feature in its own right. Von Hagen said something practical too: right now Devin can only do one pull request at a time, and he thinks there's a lot of value in having Poke orchestrate different Devin sessions and remember tasks across them. A persistent coworker, he called it. So it's part personality, part orchestration layer. If you're building an agent product, the lesson is that the interaction surface, the memory, the way it feels to work with, that's now on the roadmap next to raw capability. The robot that's pleasant to delegate to beats the robot that's marginally smarter but a pain in the neck. Ask anybody who's ever had a difficult but brilliant employee.

Alright, let me pour you the next course, and it's a bittersweet one. TechCrunch ran a piece on what it's like to be a founder under twenty right now, and I've gotta tell you, as the guy who's been watching this business a long time, parts of it made me sad and parts of it made me nod.

The lead character is Arlan Rakhmetzhanov, nineteen, from Kazakhstan. Started coding at fifteen, cold-DMed every Y Combinator founder he could find on LinkedIn until one wrote him an angel check at seventeen. His company, Nozomio, is an API index for AI agents, tool that helps agents find and use software services, and it's raised over six million bucks. And here's his quote: "I either win or lose, and a lot of young founders have the same mindset. They just want to win." He literally frames it as: build a company as valuable as Google, or fail and end up on the streets. Nineteen years old, and those are the only two outcomes he can see.

There's another founder, Pranjali Awasthi, also nineteen, dropped out of high school, then dropped out of Georgia Tech, built a startup called Slashy billed as the Cursor for emails, and she's already on to a new stealth thing. She said something that landed for me. "In 2004 you could quietly iterate for years without anyone watching. Now there's this constant ambient pressure from LinkedIn and Twitter where every raise, every milestone, every pivot is public." Build in public, sure. But also fail in public. Every stumble gets dissected by strangers in real time.

An investor in the piece, Ashley Smith at Vermilion, was honest about it. She said the market has become merciless. "It doesn't give you room to learn slowly anymore." And another founder, Aidan Guo, twenty years old, put the emotional cost right on the table: "You already have a constant fear of failure on your mind. You have to steer the ship and learn all these things as you go. And everything can always go wrong at once. And then you have all these people piling on anything you do wrong. I think people need to be more empathetic."

Now here's my kitchen-table take on this. The AI tools genuinely have democratized building. A kid in Kazakhstan can cold-DM his way into an angel check and ship a real product without ever setting foot inside a FAANG company. That's a beautiful thing, and I don't want to pretend otherwise. But the money came with strings that got a lot tighter. The article makes the point that the old forgiveness, the assumption that you'd iterate your way to product-market fit over a couple years, that's gone. Everyone's looking for the next Cursor, even though that trajectory is an outlier, not the norm. So you've got kids flush with millions expected to deliver growth in months. And that's where it gets dangerous, because the piece notes that pressure pushes some young founders toward murky ethical territory, predatory deal terms, inflated revenue numbers, content creation for social media crowding out actually writing good code.

That last part is the one I'd underline for any founder listening, young or not. When the performance of being a successful founder starts eating the hours you'd spend being a successful founder, you've got the incentives backwards. And the investor quote that closes the piece is the one worth taping to your monitor. The fundamentals of a good startup haven't changed: conviction, intellectual honesty, and obsession with the customer. As one of the founders put it, the best product that stays active and talks to customers wins. None of that has anything to do with age, and none of it has anything to do with your launch video.

Now let me connect that to something, because there's a thread here. All that founder anxiety about hitting the north-star number, all that fear of the market being merciless, that's downstream of the exact market psychology we spent time on earlier this week when we talked about whether cheap Chinese models were going to undercut everybody's revenue. The founders feel merciless pressure because the capital feels merciless pressure. It rolls downhill. And speaking of that pressure showing up in unexpected places, let's talk about crime.

Two Volkswagen engineers got charged by the Justice Department with insider trading tied to Volkswagen's joint venture with Rivian. This is the EV maker and the German automaker teaming up on electric vehicle architecture and software, a deal that started at five billion and has grown to five point eight billion, with Volkswagen now Rivian's largest shareholder. That JV was announced back in June of 2024. These two engineers, Michael Stamp and Marcus Plank, allegedly bought Rivian stock and options after they learned internally about the deal, codenamed Project Climb, but before it was public. Rivian's stock popped twenty-three percent on the announcement, they sold, and they cleared north of three hundred grand between them.

Now here's the part that's almost too perfect. Eight days before the announcement, Stamp allegedly searched "statute of limitations insider trading." And Plank's close family member searched, in German, "how is insider trading prosecuted?" Kiddos, I've said this before and I'll say it again: your search history is a confession waiting to happen. These two are looking at up to twenty-five years if convicted. Volkswagen, for its part, says the action is focused on specific individuals and doesn't involve allegations against the company. Which is the thing every company says. But the lesson for anyone building or working inside a company handling material non-public information: the tools that make you feel anonymous are the same tools that build the case against you. Google knows. Google always knows.

Let me stay on the theme of "your devices are not your friends at the worst moment," because there's a genuinely important case here for anyone who travels. TechCrunch's security editor Zack Whittaker reported on a case that's believed to be the first of its kind in the United States. The Justice Department is prosecuting an American, Atlanta resident Samuel Tunick, for allegedly wiping his phone using a "duress" password during a border search.

Here's the setup. This feature lives in GrapheneOS, a custom Android operating system that runs on a lot of Google Pixel devices. You can set a special passcode that, instead of unlocking the phone, deliberately wipes it. So when border agents at Atlanta's airport asked Tunick for his passcode, they entered it, and, according to the filing, the screen went blank, flashed a few times, and the phone restarted. Wiped. Prosecutors charged him under a federal statute that makes it unlawful to knowingly destroy or damage property to prevent authorities from seizing it.

His attorneys are fighting hard. They filed a motion to suppress, arguing the whole detention and seizure was unlawful, that he was denied access to an attorney, and that the government's stated pretext, searching for child exploitation imagery, was a cover for investigating his ties to an environmental movement, Defend the Atlanta Forest, that opposes a police training campus locals call Cop City. And the government's whole legal theory rests on that old, aggressive claim that the border isn't really US soil until you're formally admitted, so they don't need a warrant.

Now, the security experts TechCrunch talked to said they'd never seen charges brought this way before. Runa Sandvik, a digital security expert, gave the practical advice, and I want to pass it along because it's genuinely useful: the takeaway is not to have sensitive data on you when you cross certain borders in the first place. As she put it, with a little planning, you can always download the data you need once you get where you're going. The court's expected to rule on the motion to suppress later this year, and I'll tell you, builders and privacy folks should watch this one closely, because if the government wins the argument that using a built-in security feature is itself a crime, that reframes what "destroying evidence" even means in a world where your phone does that automatically.

And that lands us right at the next dish, because it's the same fight from a different angle: governments trying to control software they don't like. India moved against Jack Dorsey's Bitchat. Now Bitchat, for those who haven't seen it, is an offline, Bluetooth-powered messaging app. It runs on a mesh network, no central servers, works even when the internet's shut off. Dorsey posted on X what he said was a notice from India's Ministry of Home Affairs directing GitHub to restrict access to three Bitchat repositories within three hours.

Here's what makes this a genuinely new and, frankly, alarming legal move. The notice doesn't point to any specific illegal post or message or piece of content. It argues that the app's very architecture, its ability to work during internet shutdowns without central servers, could facilitate unlawful activity and hamper what they call "lawful interception." In other words, they're not saying "take down this bad thing." They're saying "this software shouldn't exist because of how it works."

And the context matters. This comes as India tightens internet restrictions during weeks of student-led protests in New Delhi over alleged exam paper leaks, a movement they're calling the "cockroach" movement. When authorities suspended internet service, protesters downloaded offline apps like Bitchat and Briar. And the numbers exploded. Market intelligence from Sensor Tower showed India went from about one percent of Bitchat's global downloads to about eighty-five percent in a single week. Ninety-one thousand downloads in five days, downloads jumping thirty-two-fold in a single day.

Now here's the beautiful, damning irony that the Internet Freedom Foundation pointed out. The takedown fails on its own terms. Deleting a GitHub repository doesn't delete the app from any phone that already has it, and the mesh keeps working without servers anyway. So what does the takedown actually accomplish? It prevents scrutiny of the code. That's it. And another advocate, Raman Chima, framed the deeper danger: the government isn't just targeting the service, they're trying to say that open-source development of this type of product shouldn't occur at all.

For you builders, especially anyone working in crypto, decentralized systems, privacy tech, this is the story to internalize. We are watching governments shift from "take down the illegal content" to "the tool itself is the problem." That's a categorically different threat model. If your product's value proposition is that it works when the network's down and can't be centrally controlled, you are now, by design, in the crosshairs. GitHub, for what it's worth, said the repos remained accessible in India, and they publish every takedown request they act on. But the direction of travel here is unmistakable, and it's the same current running through that duress-password case. The state wants a seat at the table inside your software, and the builders keep designing tables with no room for it.

Now let me shift gears entirely, from the ground to orbit, because there was a piece from Ars Technica by Stephen Clark about what may be the most sophisticated servicing satellite anybody's willing to talk about. I'll frame this correctly: the underlying launch happened earlier and Ars is telling the fuller story now, so this is a recent deep-dive rather than breaking-this-morning news. But it's fascinating and it matters for anybody thinking about space as a business.

The thing is called the Mission Robotic Vehicle, built by Northrop Grumman, and it launched on a Falcon 9 headed for geosynchronous orbit, twenty-two thousand miles up. What makes it special is two flexible robotic arms with seven degrees of freedom, cameras, sensors, the ability to autonomously approach another satellite, grab it, inspect it, and service it. And the key innovation over Northrop's older vehicles is that it doesn't have to commit to a single client. It carries these propulsion pods, each about the size of a dishwasher, and it can install them onto multiple aging satellites to extend their lives, like giving an old spacecraft a jetpack.

The economics here are what I want you to hear. The whole business is a bet against what they call the "launch-and-abandon" model. Right now, a multi-billion-dollar military or communications satellite runs out of fuel or breaks a part, and it's just dead. Garbage in orbit. The pitch is: what if you could send a repair truck up there instead? DARPA spent about four hundred and twenty million on the robotics payload, Northrop put in hundreds of millions more, and the government's clever angle is that Northrop owns and operates the thing for ten to thirteen years, so Uncle Sam gets the capability without paying for the long-term care and feeding. They just contract for services as needed.

The program manager, Jim Shoemaker, described why this is so hard, and it's a nice little physics lesson. "Inertia operates differently in space. When you move the arm one direction, the entire satellite wants to rotate the opposite direction." And then when you grab onto another satellite, you've doubled your mass, so the control system has to adapt to a whole new configuration. As he put it, "These are things that tend to be really hard." And you can't joystick it from the ground, the time delay is too long, so it has to execute a mission script or run autonomously.

Now, why should a founder care about robot arms in space? Two reasons. One, there's a real emerging market here, in-orbit servicing, refueling, orbital logistics, and it's got national security money flowing into it. The article notes China's been demonstrating similar capabilities, with US Space Command openly worried that these dual-use servicers could be turned to offensive purposes against satellites. So there's a strategic race dimension. And two, it's a clean example of a public-private structure where the government de-risks the R&D and industry owns the operations. If you're building anything hardware-heavy and capital-intensive, that DARPA-to-industry handoff model is worth studying, because it's how expensive, decade-long bets actually get financed.

While we're up there, quick lightning round from the same Rocket Report, and I'll note this collects a few developments from recent weeks rather than one thing that happened overnight. India's first fully commercial rocket, Skyroot's Vikram-1, reached orbit, a genuine milestone for private launch outside the usual players. The US Space Force tripled the maximum value of one of its launch contracts to seventeen billion, pushing its combined launch spending past thirty billion, which tells you demand for putting military hardware in orbit is only going one direction. And Russia's answer to the Falcon 9, the Amur rocket, announced in 2020 for a 2026 debut, has now slipped to 2031. Ars put it perfectly: year for year, its launch date moved five years further away. So basically it's standing still while time passes. Some projects, kiddos, are permanently one financing round from launch.

Alright, let me bring us back down to earth for the last real story, and this one's about who controls what the public gets to see during a war. Ars Technica's Jeremy Hsu reported that the European Union has agreed to a US request to delay the release of satellite images of the Iran war region. Specifically, they're now holding back images of the shipping lanes near the Strait of Hormuz by twenty-four hours.

Here's the mechanism and why it's a big deal. Europe runs a program called Copernicus, an Earth-observation system with a free-and-open data policy. Anybody can pull the images. Journalists, researchers, open-source intelligence groups, they've relied on it heavily, especially because US commercial firms like Planet Labs and Vantor have already restricted their imagery of the Iran conflict at the US government's request. So Copernicus became the fallback for independent transparency. Bellingcat built a tool on it to map war damage. The New York Times and Washington Post used it to investigate damage to US military bases from Iranian strikes.

And now, per a decision by the Council of the EU that Space News got hold of, that fallback has a twenty-four-hour delay bolted onto it. The US first made the request back in May. And it comes at a time when the Pentagon has become notably less willing to share information about US casualties or damage to its assets. The war, which the piece notes has already killed thousands and triggered a global energy shock through the Strait of Hormuz, has already cost the US more than thirty-seven billion, with the defense secretary requesting another seventy billion on top.

Now, a twenty-four-hour delay is less onerous than a full blackout, and journalists can still monitor the war. But the principle here is the one to sit with. This is the free-and-open data policy of a civilian European science program getting bent to a wartime information-management request from a foreign government. It's one more data point in a week absolutely soaked in the same theme: the tools that provide independent transparency, satellite imagery, offline messaging, encrypted phones, are all getting squeezed at exactly the moment people most need them. Whether it's Bitchat during an internet shutdown, or a phone at the border, or a satellite image of a war zone, the through-line is friction being deliberately added between the public and the truth.

Before I let you go, one quick note on the mood music. Sam Altman posted a short one on X, saying he wants the US to win in AI in both open-source and proprietary models, and that he's glad to see something on that front. I won't over-read a one-liner, and I'm not going to speculate about what specifically he was reacting to since the post just points to a link. But I'll flag it because it fits the larger policy fight we've been tracking, the one about open weights, distillation, and how much America's AI lead depends on staying open versus locking things down. When the CEO of the most proprietary shop in town is out here waving the flag for open source too, that tells you the open-versus-closed debate has gotten political enough that everybody wants to be seen on the right side of "American AI winning." File it under: watch this space.

So let me tie the whole plate together before I turn off the mic. The big builder takeaway today is Opus 5: smarter judgment, half the price of the flagship, fewer wasted tokens, and a deliberate wall between finding vulnerabilities and exploiting them. That's the practical one, go kick the tires. But the theme that kept surfacing across everything else was control and transparency, who gets to see, who gets to build, who gets to speak when the network goes dark. From a kid in Kazakhstan building in public under a microscope, to a guy wiping his phone at the border, to India trying to erase software off GitHub, to satellite pictures of a war held back a day. Different stories, same nerve. And as the tools get more powerful, that nerve is only going to get more raw.

Chew on that this weekend. That's the menu for today. I'm Tony DeLuca, this has been Barely Possible, and I'll be right back here next time with another fresh plate. Take care of each other out there.