UpNext AI

The U.S. Department of Defense is reportedly testing competing frontier AI models as it evaluates alternatives to Anthropic’s Claude. Bloomberg reports that a group of Pentagon “power users” is comparing models in real operational workflows, highlighting a broader shift from benchmark-driven competition to real-world evaluation focused on reliability, mission fit, security, and deployment requirements. For AI vendors, winning enterprise and government adoption increasingly depends on performance in production environments rather than leaderboard rankings alone.  
Meanwhile, agent infrastructure startup Daytona argues that AI agents need something beyond model APIs: actual computers to operate. In a Latent Space interview, CEO Ivan Burazin said the company has experienced rapid growth as coding agents, evaluation systems, and reinforcement learning workloads increasingly require isolated, stateful environments. The broader trend is clear: a new infrastructure layer is emerging between foundation models and applications, designed specifically for autonomous agents and long-running workflows.
In research, we examine a study in Scientific Reports exploring AI-based safety forecasting for extreme cold exposure. Researchers developed an LSTM model to predict toe skin temperature in mountaineering conditions and introduced a metric called Duration of Safe Exposure. Rather than optimizing only for prediction accuracy, the system was designed to minimize dangerous forecasting errors where risk could be underestimated. The work highlights a growing theme across applied AI: success is increasingly measured by safety and decision quality, not just average model performance.
In the headlines: President Trump delays an executive order that would have expanded government evaluation of advanced AI models before release, Amazon Bedrock adds request-level AI usage attribution for enterprise cost tracking and governance, Google continues rolling out Gemini, Search, and smart-glasses initiatives following I/O 2026, and Anker introduces its first earbuds powered by an in-house AI audio chip for enhanced noise reduction and voice processing.

Sources
Bloomberg – Pentagon tests rival AI models as alternatives to Anthropic
 https://www.bloomberg.com/news/articles/2026-05-21/pentagon-tests-rival-ai-models-in-race-to-replace-anthropic
Latent Space – Giving Agents Computers (Ivan Burazin, Daytona)
 https://www.latent.space/p/daytona
Nature Scientific Reports – LSTM-based safety-oriented prediction of toe skin temperature in extreme cold conditions
 https://www.nature.com/articles/s41598-026-52990-x
TechCrunch – Trump delays AI security executive order
 https://techcrunch.com/2026/05/21/trump-delays-ai-security-executive-order-i-dont-want-to-get-in-the-way-of-that-leading/
AWS – Amazon Bedrock request-level usage attribution
 https://aws.amazon.com/about-aws/whats-new/2026/05/amazon-bedrock-request-level-usage-attribution/
WIRED – Everything announced at Google I/O 2026
 https://www.wired.com/story/everything-google-announced-at-google-io-2026/
The Verge – Anker’s AI-powered Liberty 5 Pro earbuds
 https://www.theverge.com/tech/934621/anker-liberty-5-pro-max-wireless-headphones-earbuds-ai-thus-chip

What is UpNext AI?

Daily AI news and research, distilled. UpNext AI breaks down the most important developments in artificial intelligence—from major industry moves to cutting-edge papers.

Welcome to the UpNext AI podcast. It's Friday, May 22nd, 2026, and here's what matters in AI today.\n\nBloomberg reports the Pentagon is testing competing AI models with 25 of the department’s so-called power users, as the US military looks for alternatives to Anthropic’s Claude.\n\nThat makes this more than a routine procurement story. It’s a snapshot of how frontier models are starting to be compared in live, high-stakes operating environments, not just on benchmarks or internal demos. If those evaluations spread across defense and other public-sector buyers, model competition starts to look less like brand preference and more like an ongoing bake-off tied to mission fit, reliability, and trust.\n\nThe detail Bloomberg gives us is narrow but important: this testing is being driven by user preference among a specific group of heavy internal users. So the immediate signal here is not that one company has won or lost, but that the Pentagon is actively widening the field.\n\nAnd that matters for the rest of the market too. When a buyer this large starts formalizing side-by-side model evaluation, it can influence how vendors position themselves on security, deployment control, and specialized performance.\n\nFrom there, let’s move from model choice to the infrastructure agents run on.\n\nIn a conversation published by Latent Space, Daytona CEO Ivan Burazin describes a fast-growing market for what he calls computers for AI agents. The company says one customer runs about 850,000 sandboxes a day, that Daytona has seen 74 percent month-over-month growth, and that reinforcement-learning and evaluation workloads have jumped from effectively zero to roughly half of usage in just a few months.\n\nThe core idea is simple: a lot of agents don’t just need to generate code or text. They need an environment they can actually operate in — something stateful, isolated, fast to start, and flexible enough to handle messy real workflows.\n\nAccording to the interview, Daytona is betting that this category looks less like disposable code execution and more like API-accessible compute environments that can scale very quickly. The company also claims it can spin up a single sandbox in around 60 milliseconds and 50,000 sandboxes in about 75 seconds.\n\nBecause this comes through a company-centered interview, the clean read is not that Daytona has single-handedly defined the category. It’s that the agent market is creating real demand for infrastructure between the model and the application layer — especially for evals, coding agents, and other workloads that need an actual machine to work on.\n\nFor today’s research note, a paper in Scientific Reports looks at a very practical safety problem: how long mountaineering footwear can keep someone safe in extreme cold.\n\nThe researchers trained an LSTM model to predict big-toe skin temperature over time under different combinations of footwear insulation, ambient temperature, and physical activity. They also introduced a metric called Duration of Safe Exposure, or DSE — basically, how long it takes before toe temperature falls to a conservative danger threshold of 15 degrees Celsius.\n\nWhat makes this interesting is that the model was not just graded on average prediction error. The study also focused on whether the system correctly predicted the threshold-crossing event, and whether it made unsafe mistakes by overestimating how long conditions would stay safe.\n\nIn the reported setup, unsafe predictions were rare, and the researchers say most estimates were either conservative or within tolerance. They also linked the trained network with a thermoregulation model to explore virtual scenarios.\n\nThe catch is that this was a controlled experimental setup with 96 time series, and the participants were exclusively young male subjects, so this is not a universal cold-safety engine.\n\nBottom line: the paper shows how AI can be tuned for safety-oriented forecasting, where the key question is not just what happens on average, but whether the model misses the moment risk becomes real.\n\n...Are you building apps with voice? Elevate your app's voice capabilities with ElevenLabs. Their API is a game changer for embedding dynamic, responsive voice interactions in your applications, providing unprecedented realism, flexibility and latency. In fact, you're listening to one of their voices - right - now. If you are a developer looking to elevate user experience with natural voice interfaces, this is your solution. Visit up next dot fm slash eleven to check out their latest offerings. ...\n\nFirst, TechCrunch reports President Trump delayed signing an executive order that would have allowed the government to evaluate AI models before release. Trump said he was unhappy with parts of the language, and the reporting says one sticking point was a proposal that companies share advanced models with the government between 14 and 90 days ahead of launch.\n\nNext, Amazon says Bedrock now supports request-level usage attribution on the InvokeModel and InvokeModelWithResponseStream APIs. That means customers can tag inference usage to specific teams, applications, environments, and experiments, giving enterprises a more detailed way to track AI consumption and spend.\n\nAlso, Google’s I/O announcements are still landing across the industry. Wired’s broad recap frames the event around Gemini updates, a revamped search experience, AI agents across products, and smart glasses expected this fall. We covered Google’s search shift earlier this week, but the bigger takeaway is that Google is continuing to push AI as a cross-product layer rather than a single feature set.\n\nAnd finally, The Verge reports Anker’s new Soundcore Liberty 5 Pro earbuds are the company’s first to use its Thus AI audio chip for stronger noise reduction and clearer calls in loud environments. The earbuds start at $169.99, and a Max version adds AI note-taking through the charging case.\n\nBefore we wrap up, a quick note: this podcast is generated with the assistance of AI and is intended for informational purposes only. All referenced articles, research, and commentary remain the property of their original authors and publishers.\n\nIf you enjoyed this episode, don't forget to subscribe, rate, and leave us a review! And that's your briefing for today. Full source links are in the episode notes. We'll be off Monday for the US holiday, but we'll be back Tuesday with what's up next!