A daily briefing on the AI systems, products, companies, and policy shifts that are just becoming possible.
Want a podcast for your own topics? Join early access: https://www.barelypossible.to/waitlist/?source_path=public_feed&feed_source=rss
Okay kiddos, I'm your boy Tony DeLuca, and we've got a fresh plate of tech, money, and a little bit of nonsense to work through today, so grab your coffee and let's dig in.
I want to start today with a story that's easy to laugh at and then get very serious about. There's a company called Kalshi. Prediction market. They'll tell you it's not gambling, it's "forecasting the future," it's contracts and information, more like a soybean futures contract than a pull on the one-armed bandit. That's their line. And states like New York looked at it, squinted, and said, that sure looks like betting to me, and tried to regulate it under gambling laws. So Kalshi went and got itself under the Commodity Futures Trading Commission, the CFTC, for federal cover. Which means the CFTC is now out there suing states — Kentucky, Minnesota, Illinois, Rhode Island — to knock down their laws in favor of one national standard the CFTC controls.
Now here's where it gets good, and this is a report that surfaced this week, though the underlying events ran from late last year into the beginning of this year. According to reporting from NPR and ABC, President Trump's teleprompter operator — a guy named Gabriel Perez — allegedly made a hundred grand betting on what words would come out of Trump's mouth during speeches. There's a whole corner of Kalshi called a "mention market." You can bet on what Domino's will say during their next earnings call. I'm not making that up — the smart money apparently thinks "Parmesan" and "DomOS" get mentioned. And Perez, the guy who has the final eyes on nearly all the prepared remarks and takes last-minute edits from Trump himself, was reportedly sitting there adjusting his bets mid-speech. When Trump skipped over a section that had a word Perez had money on, he'd back out of the bet. Live. During the speech. On his phone.
Picture that. The teleprompter guy for one of the most powerful people on earth, tapping away at his phone to protect his position. Kalshi flagged the unusual activity, found out it was a federal employee, froze his funds, and sent it to the CFTC. They're in settlement talks. Will he be prosecuted? Nope — the feds in Manhattan declined to open a criminal case. He just doesn't work at the White House anymore.
Now why does a founder-and-builder crowd care about a teleprompter guy's side hustle? Because this is the tell on where a whole category is heading. Prediction markets spent heavily advertising during the World Cup, and they're extending "forecasting" from sports into drug trials, flight cancellations, and the words people say in speeches. As McKay Coppins laid out in a recent Atlantic piece, roughly half of men ages eighteen to forty-nine already have an active online sportsbook account. When you turn every event into a tradable contract and you wrap it in the language of markets to dodge gambling law, you get insider trading stories that are equal parts absurd and corrosive. The teleprompter guy is the punchline. The structure underneath him is the actual story, and if you're building anything in the crypto or fintech or prediction space, that regulatory arbitrage — CFTC shield versus state gambling law — is the fight to watch.
Alright, that's the appetizer. Let me line up the menu, because there's a real spread today: an AWS billing bug with a comically large number on it, a nuclear startup raising money on the back of the AI buildout, humanoid robots setting up shop in Tesla's backyard, the EU putting the screws to Google, and OpenAI's CFO trying to tell you how to measure whether any of this AI stuff is actually paying off. And that last one is where we're going to spend our real time today.
Now let's get into a story that made a lot of cloud customers spit out their coffee Friday morning. Amazon confirmed it was fixing a bug in the AWS billing portal that told some customers they owed millions — and in at least one case, according to screenshots posted on Reddit, close to two and a half billion dollars — for cloud services they never used. Two and a half billion. For one month. Amazon said the inaccurate billing data started late Thursday, and here's the part I love: they tried to roll back a recent change to their billing computation subsystem, and the rollback didn't fix it. The good news for anybody who had a heart attack Friday — the estimates don't reflect actual usage, you're off the hook.
But let me sit on this for a second, because I think there's a real lesson buried in a dumb headline. If you're a builder, your cloud bill is not just an operational number, it's a trust number. The whole promise of the cloud is that you don't have to think about the plumbing — you swipe the card, it just works, the meter's honest. And when the meter goes haywire and tells you that you owe more than the GDP of a small country, even for a few hours, it pokes a hole in that. Amazon wouldn't say whether any accounts got suspended or paused because of it. That's the question that actually matters for a founder — not the funny number, but whether an automated system somewhere downstream acted on that funny number before a human caught it. That's the recurring theme of this whole era, isn't it? The systems are fast and mostly right, and then every so often they're spectacularly, automatically wrong, and the only thing standing between you and chaos is whether a person was watching. Keep your billing alerts sane, folks, and maybe don't wire your account suspension logic directly to the estimate.
Speaking of the cloud eating the world's electricity, let's talk about money chasing power. There's a nuclear startup called Valar Atomics, three years old, and it's in talks to raise a new round at about a six billion dollar valuation, with Sequoia expected to lead. The Information reported it first — a one billion dollar equity round, though part of that was raised earlier at a lower mark. And here's a detail worth chewing on for the founders listening: Valar had previously raised four hundred fifty million at a two billion dollar valuation, per a Bloomberg report back in March. So in the space of a few months, on paper, they triple.
Now the piece makes an important point about how these deals are getting structured — capital coming in multiple installments at different valuations, sometimes at different times, which creates the illusion that everybody paid the same price when they didn't. That's becoming standard in this AI-fueled fundraising environment, and if you're trying to benchmark one red-hot startup against another, that distinction matters more than ever. Two companies can both say "six billion dollar round" and mean completely different things about who actually paid what.
What's Valar building? Small modular reactors — miniaturized, factory-built power plants meant to be cheaper and faster to deploy than the traditional giants. Helium-cooled, high-temperature gas reactor. Earlier this month they showed their reactor providing a small amount of power to an Nvidia AI chip, and announced a partnership with Nvidia to explore nuclear for future AI data centers. Their backers include Palmer Luckey from Anduril and Shyam Sankar from Palantir. The founder, Isaiah Taylor, dropped out of high school at sixteen, he's twenty-seven now, and he likes to mention his great-grandfather worked on the Manhattan Project.
Here's my skeptical-neighbor take. The demand crunch is real — data center electricity needs are projected to grow sharply, utilities are years behind on capacity, and that vacuum has turned nuclear, long buried under cost overruns and regulatory quicksand, into one of the hottest corners of the AI infrastructure boom. But "showed its reactor powered a chip" is a proof of concept, not a power plant. SMRs are theoretically cheaper to build, but the technology is still nascent, and nobody can tell you when it deploys at industrial scale. And notice — Valar's already suing its own regulator, the Nuclear Regulatory Commission, arguing the agency wrongly applies the same lengthy licensing process to small test reactors that it uses for full commercial plants. That litigation keeps getting paused, which hints at a settlement in the works. So what you've got is a company valued on a demand curve and a regulatory bet, not on kilowatt-hours delivered. That can absolutely work out. Just know which one you're buying when the round is this frothy.
Now let's shift from the power grid to the factory floor, and I'll connect this to something we covered yesterday. Yesterday we talked about the Hyundai union walkout in Ulsan — the car industry's first factory stoppage over humanoid robots, workers scared of Boston Dynamics Atlas units. Today's story is the other side of that same coin, the company side, and it's a fresh one out of TechCrunch. Agility Robotics is opening a sixty-thousand-square-foot facility to train its humanoid robots in Fremont, California — just up the highway from where Tesla is expected to start building its Optimus robots this year.
And I appreciated the tone from Agility's CEO Peggy Johnson. She said, and I'm quoting, "It's great to have Tesla in the same area as us, because really, for a long time Agility was out there alone." That's a company that's confident. And they have a right to be, because unlike a lot of the humanoid hype, Agility has a robot named Digit that's actually generating revenue right now — carrying totes and bins in warehouses for customers like Amazon, GXO, Schaeffler, and Toyota Motor Manufacturing Canada. They say they've got three hundred million in contract orders. Digits have moved a hundred thousand totes at one GXO facility.
Here's the part that matters for how you think about building with AI, and it's a beautiful bit of engineering judgment from co-founder Damion Shelton. He said, when you think about self-driving cars, you really don't want the anti-lock brake controller under AI control. The analog with humanoids is that all the safety stuff needs to go through a path that's not generative AI. Quote: "You don't want to get creative with your safety stack." That's the whole ballgame, folks. That's a grownup talking. The AI does the flexible, imaginative part — figuring out the thousand things a robot could be asked to do, because as Shelton put it, the number of things you can imagine a robot doing is far larger than the number of engineers who can program robots, and generative AI answers that question. But the brakes, the safety, the stuff that kills people if it hallucinates? That runs on deterministic, boring, non-creative code. Digit still works in a human-free zone today; version five, coming this fall, will finally be able to sense humans and work alongside them.
And Agility's going public through a reverse merger, expected to be the first pure-play humanoid robot company on the public markets later this year. So while everybody's dreaming about robots in your living room, Agility's co-founder Jonathan Hurst laid out the actual roadmap: start with bins and totes, then picking and kitting, then the really hard stuff like cardboard and loading tractor trailers. And then, he says, half-joking, "now we're at a hundred million robots, a trillion-dollar company." Bins first. Trillion dollars later. That's a company that knows the difference between a demo and a deployment — which, if you've been listening this week, is the thing I keep coming back to.
Let me stay on the theme of who gets to decide how things work, and move to Europe. This is a recent report, from mid-July — the European Commission made it official: under the Digital Markets Act, they're forcing Google to share search data with competitors and open up AI access on Android. Two big pieces. On Android, Google has to open the system to competing AI assistants — right now Gemini gets the preferential treatment, it's preloaded, it answers to "Hey Google," it's got deep system and app automation access, and third-party AI assistants can't get near that. The Commission says that limits competitors and makes them less attractive to the sixty percent of EU users on Android. On the search side, Google has to share search data with rival search providers, transparently and for a reasonable fee, and — this is the interesting one for builders — Google has to treat AI chatbots as search services for the purposes of data sharing.
Google's not happy. Kent Walker, their president of global affairs, said the decisions "risk undermining vital privacy and security guardrails for millions of Europeans." Which — look, I'm skeptical of hype in both directions. When a dominant company tells you that letting competitors in will hurt your privacy, you're allowed to notice that the thing being protected also happens to be their moat. Google has to be ready to share search data by January 2027, and the Android AI opening has to be done by July 2027. For anybody building an AI assistant or an alternative search product, that mandated access to Google's search metrics is potentially a very big deal — it's the kind of thing that turns "you can't possibly compete with Google's data advantage" from a law of physics into a regulatory line item.
Alright. Now let's dig into the story I think matters most for you today, because it's the one that's aimed squarely at the question every founder and every enterprise buyer is quietly sweating right now: is any of this AI spend actually working?
On Thursday, OpenAI put out a piece from its CFO, Sarah Friar, and the framing is right there in the title — an essay called A scorecard for the AI age. The pitch is a practical AI scorecard to measure return on investment through four things: useful work, cost per successful task, dependability, and return on compute.
Now, I want to be honest with you about what this is and what it isn't. This is the CFO of the company selling you the tokens telling you how to measure whether the tokens are worth it. So you read it with one eyebrow up. But — and this is the important but — the fact that OpenAI's own CFO is now leading with cost per successful task instead of raw capability, instead of benchmark scores, instead of "look how smart it is," that itself is the news. That's a tell about where the whole industry's conversation has moved. We spent two years talking about how good the models are. Now the money people are talking about whether the good models pay for themselves.
Let me connect this to something small but sharp from earlier in the week. Simon Willison — a guy who watches this stuff about as closely as anyone — put out a note that stuck with me. He said, and I'm quoting, "Right now it feels like the single biggest competitive advantage an AI lab could have is making it abundantly clear whether and how they will train models on your data." And then the kicker: "I pay pretty close attention to this and I couldn't confidently summarize the policies for ANY of the lead labs."
Sit with that. Somebody who does this for a living cannot confidently tell you the data-training policy of a single leading lab. So on one hand you've got OpenAI's CFO handing you a clean four-box scorecard — useful work, cost per successful task, dependability, return on compute. Very tidy. Very CFO. And on the other hand you've got one of the sharpest independent observers in the field saying the most basic trust question — what do you do with my data — is a fog for every one of them. Those two things belong in the same conversation, because a scorecard that measures cost per successful task but can't answer "and where does my proprietary data end up" is measuring the easy half of the equation and skipping the half that actually keeps enterprise lawyers up at night.
Here's why I think the Friar scorecard is genuinely useful anyway, even coming from an interested party. Cost per successful task is the right unit. Not cost per token. Not cost per query. Cost per successful task. Because the dirty secret of this agentic era — and we've been living in it all week on this show, the file-deleting, the credential-hunting, the overly-agentic behavior — is that a lot of the tokens you're paying for are the model spinning, retrying, going down a wrong path, burning compute on work that doesn't land. If your metric is cost per token, an efficient model and a model that flails for twenty minutes look similar on the invoice. If your metric is cost per successful task, the flailing model gets exposed. That's the number that actually tells you whether you've got a product or a science experiment.
And "return on compute" is the phrase I'd tattoo on the wall of every startup burning venture money on inference right now. Because remember what Sriram Krishnan flagged, and what we've been tracking — companies have shifted from "how do we get our people to use more tokens" to being genuinely uncomfortable with their token costs ballooning with no clear line to revenue. A year ago the anxiety was adoption. Now the anxiety is the bill. The scorecard is OpenAI reading that room and trying to give you a vocabulary to justify the spend — which, cynically, is also a vocabulary to justify keeping the spend going.
So here's my plain-language advice to the builder listening. Take the Friar scorecard, it's a good skeleton — cost per successful task, dependability, return on compute. Steal it. But bolt on the thing the vendor's scorecard conveniently leaves off: data governance clarity. Before you scale spend on any lab, you should be able to answer, in one sentence, whether and how they train on your data — and if you can't, that's not a footnote, that's a line item on your risk register. Willison can't answer it for the lead labs. Make sure you can answer it for yours before your usage — and your exposure — quietly compounds.
And notice how this connects to a couple of the OpenAI-side items floating around this week — the case studies about Cars24 running voice and chat agents at over a million conversation minutes a month and recovering twelve percent of lost leads, and the company's push on making ChatGPT safer for teens with parental controls and age-appropriate protections. Those aren't random. That's a company assembling a story: here's the ROI proof, here's the scorecard to measure it, here's the safety posture so your compliance people don't panic. It's a coordinated pitch to the enterprise buyer. As a builder, you don't have to reject the pitch — you just have to read it as a pitch, run your own cost-per-successful-task numbers, and get your own answer on the data question.
Alright, let me bring us home with a couple of quicker ones, because they round out the picture.
There's a resurfaced story worth a brief callback — San Francisco's city attorney David Chiu ordered Apple and Google to purge so-called "nudify" apps from their app stores, apps that digitally strip clothing off real people's photos. California law criminalizes knowingly facilitating non-consensual deepfake pornography, and Chiu's office estimates Apple and Google likely made millions in fees off these things. Both companies say they've suspended the flagged apps. The reason I mention it in a builder's episode: the researchers found that some of these apps evade app-store removal by only advertising the innocent face-swap feature while hiding the harmful capability — in one study, seventy percent of "face-swap" apps tested could actually be used to nudify. If you run a platform, that's the enforcement problem of the decade: the harm is in what the tool can do, not what its listing says it does. And it's sitting right next to the unresolved Grok question — whether the company that makes the model, or just the user who prompts it, carries the liability when the outputs are illegal.
And one for the road, because I can't resist. There's a fun piece on the Russian answer to the Falcon 9 — a rocket called Amur that Roscosmos announced back in 2020, promising a 2026 debut. Well, it's 2026. And this week a senior Russian official said the current focus is on building a demonstrator, a hopper vehicle, with tests beginning in 2028 — and a placard at a security forum earlier this year now lists flight tests in 2031. So the launch date has slid five years to the right in the six years since they announced it. There's a founder lesson buried in the rocket story too: the gap between the press release date and the actual date tells you almost everything about whether a thing is real. SpaceX landed its first orbital rocket a decade ago. Amur is a schematic and a slipping calendar. Slope, not the y-intercept, as one guy said this week about something else entirely — but it applies here just fine.
So that's the plate today. A teleprompter guy betting on Trump's vocabulary, a two-and-a-half-billion-dollar phantom AWS bill, a nuclear startup tripling its paper value on a demand curve, robots learning to move totes in Tesla's backyard, Europe prying open Google's search data, and the AI industry quietly switching the conversation from "how smart is it" to "does it pay for itself." Track that last one hardest — cost per successful task, and know where your data goes.
That's the show, kiddos. This is Tony DeLuca, telling you to keep one eyebrow up and both hands on your billing dashboard. Catch you next time on Barely Possible.