Daily AI news and research, distilled. UpNext AI breaks down the most important developments in artificial intelligence—from major industry moves to cutting-edge papers.
Welcome to the UpNext AI podcast. It's Friday, May 1st, 2026, and here's what matters in AI today.
OpenAI says it is accelerating the compute buildout for what it calls the Intelligence Age.
In a company post, OpenAI said Stargate is its long-term effort to build the compute foundation needed to deliver AGI broadly and reliably, and that it has already passed its earlier milestone of securing 10 gigawatts of AI infrastructure in the U.S. by 2029. OpenAI says more than 3 gigawatts were added in just the last 90 days.
The company’s argument is straightforward: demand for AI from consumers, businesses, developers, and governments is rising fast, and the bottleneck is increasingly physical infrastructure. In OpenAI’s telling, more compute means better models, more reliable serving, lower costs over time, and wider access.
The post also gives a clearer picture of how OpenAI wants this buildout to be understood. It emphasizes that no single company can do this alone, and frames Stargate as an ecosystem effort spanning cloud infrastructure, chipmakers, energy providers, construction firms, investors, skilled trades, and local communities.
OpenAI also points to Abilene, Texas as a flagship Stargate site. According to the company, GPT-5.5 was trained there on Oracle Cloud Infrastructure using NVIDIA GB200 systems. OpenAI says the site uses closed-loop cooling, and it presents that as part of a broader case that large-scale AI infrastructure can be built quickly while still addressing water use, workforce needs, and local economic benefits.
The practical takeaway here is that frontier AI competition is no longer just about model releases. It is also about who can secure power, land, chips, data center capacity, and construction at national scale. And on that front, OpenAI is signaling that infrastructure itself is now a product strategy.
Staying with OpenAI, but shifting from scale to access control: TechCrunch reports the company will begin rolling out GPT-5.5 Cyber only to what OpenAI calls critical cyber defenders at first.
According to the report, OpenAI has an application process for users to submit their credentials and intended use. The tool is described as supporting tasks like penetration testing, vulnerability identification and exploitation, and malware reverse engineering.
That makes this notable for two reasons. First, it is an explicit limited release of a high-risk cybersecurity capability. And second, TechCrunch frames it as a reversal in tone after Sam Altman had criticized Anthropic for restricting access to its own cybersecurity tool, Mythos.
The larger point is that frontier labs are converging on the same basic reality: some model capabilities may be useful for defenders, but they also create real misuse risk. So even when companies argue publicly about openness, in practice they are still putting gates around the sharpest tools.
OpenAI says, according to TechCrunch, that it wants to make Cyber more widely available over time, including by consulting with the U.S. government and identifying more users with legitimate cybersecurity credentials. For now, though, this is a restricted rollout, not a broad product launch.
For today’s research note, a paper on a problem that sounds narrow but matters a lot in production: evaluating text-to-SQL systems.
The paper is titled Agent-Agnostic Evaluation of SQL Accuracy in Production Text-to-SQL Systems. The core issue is simple. In demos and benchmarks, text-to-SQL systems are often graded against a known correct query and a known database schema. But in a live product, you often do not have that luxury. You may have the user’s question and the SQL your system generated, but not a neat ground-truth answer waiting in the background.
The researchers propose what they call STEF, a schema-agnostic text-to-SQL evaluation framework. Instead of depending on a reference query or full schema-aware grading, the system works from natural language inputs and the generated SQL. It tries to extract the intended meaning from the user’s request, compare that against what the SQL appears to do, and then produce an interpretable score from zero to one hundred.
The paper says that score combines things like filter alignment, a semantic verdict, and evaluator confidence. It also highlights a few practical touches meant for real production environments, including tolerance for common SQL variations like GROUP BY differences, ORDER BY defaults, and LIMIT heuristics.
Why does that matter? Because one of the hardest problems in enterprise AI is not generating something plausible. It is knowing, continuously, whether the system is still doing the right thing after deployment. If a text-to-SQL tool quietly starts drifting, teams need a monitoring layer that does not depend on perfect labels.
Bottom line: this paper is aiming at a very real operations gap in AI products, moving evaluation from benchmark theater toward something that could actually help monitor live systems.
...Are you building apps with voice? Elevate your app's voice capabilities with ElevenLabs. Their API is a game changer for embedding dynamic, responsive voice interactions in your applications, providing unprecedented realism, flexibility and latency. In fact, you're listening to one of their voices - right - now. If you are a developer looking to elevate user experience with natural voice interfaces, this is your solution. Visit up next dot fm slash eleven to check out their latest offerings. ...
And one quick headline before we go: Google and Kaggle are bringing back their free five-day AI Agents Intensive course.
Google says the online course will run June 15th through 19th, with updated content, new speakers, and a hands-on capstone project. The focus is on building AI agents, including what Google calls vibe coding workflows, where natural language becomes a primary interface for programming and tool use.
If you want a structured, no-cost way to get sharper on agent workflows, this one is worth a look.
Before we wrap up, a quick note: this podcast is generated with the assistance of AI and is intended for informational purposes only. All referenced articles, research, and commentary remain the property of their original authors and publishers.
And that's your briefing for today. Full source links are in the episode notes, and we'll be back Monday with what's up next!