Daily AI news and research, distilled. UpNext AI breaks down the most important developments in artificial intelligence—from major industry moves to cutting-edge papers.
Welcome to the UpNext AI podcast. It's Tuesday, May 12th, 2026, and here's what matters in AI today.\n\nFirst up, General Motors is reportedly making a very concrete kind of AI transition. TechCrunch reports GM laid off more than 10 percent of its IT department, about 600 salaried employees, while continuing to hire for different skill sets centered on AI.\n\nAccording to the reporting, the roles GM now wants include AI-native development, data engineering and analytics, cloud-based engineering, agent and model development, prompt engineering, and new AI workflows. In other words, this is not just a company handing existing teams a chatbot license and calling it transformation. It’s a workforce reshaping around building with AI from the ground up.\n\nGM confirmed the layoffs to TechCrunch and said it is transforming its IT organization to better position the company for the future. TechCrunch also reports the company has been shifting resources toward high-priority initiatives, including AI, over the last 18 months.\n\nThe reason this matters goes beyond one automaker. This is one of the clearest examples of what enterprise AI adoption can look like when it moves from experimentation into org charts, hiring plans, and layoffs. The AI story here is not abstract productivity. It’s a skills swap.\n\nFrom workforce restructuring to interface design: Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, says it wants to build AI that can listen while it talks.\n\nTechCrunch reports the company announced what it calls interaction models. The core idea is full duplex conversation: instead of the usual pattern where you speak, then the model speaks, this system is designed to process your input and generate a response at the same time.\n\nThe company says its model, TML-Interaction-Small, responds in 0.40 seconds, which TechCrunch describes as roughly the speed of natural human conversation. That makes this less about raw model intelligence and more about changing the feel of AI itself, from a text-message rhythm to something closer to a phone call.\n\nThe important caveat is that this is still a research preview, not a public product. TechCrunch says a limited research preview is expected in the next few months, with a wider release planned later this year.\n\nIf this works outside the lab, it could matter a lot for voice assistants, live copilots, tutoring, customer support, and any setting where waiting your turn makes AI feel unnatural. The bet here is that conversation quality is not just about what the model says, but when and how it says it.\n\nNow to the research note.\n\nA new paper on arXiv is called WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation. And the plain-English idea is simple: a lot of agent benchmarks still make agents look better than they really are because the tests are too clean, too short, and too artificial.\n\nThe researchers focus on agents powered by language and vision-language models working through command-line interface harnesses. Their argument is that if you want to know whether an agent can actually do useful work on a user’s behalf, you have to test it in something closer to real deployment conditions, not just a tidy sandbox.\n\nIn the paper text, WildClawBench is described as a native-runtime benchmark with 60 human-authored, bilingual, multimodal tasks spanning six categories. The average task takes roughly 8 minutes of wall-clock time and more than 20 tool calls. The benchmark runs inside reproducible Docker containers and uses real agent harnesses and real tools rather than mock services.\n\nEvaluation is hybrid: deterministic checks, environment-state auditing of side effects, and an LLM or VLM judge for semantic verification. Across 19 frontier models, the paper says the best result reached 62.2 percent overall, and every other model stayed below 60 percent. It also says switching the harness alone could move the same model by as much as 18 points.\n\nThat last result is especially useful. It suggests agent performance is not just about the model. The surrounding runtime and tooling stack can materially change the outcome.\n\nBottom line: if you want a realistic picture of whether agents are ready for real work, you need longer, messier tests in real environments, and on this benchmark, today’s frontier systems still look far from solved.\n\n...Are you building apps with voice? Elevate your app's voice capabilities with ElevenLabs. Their API is a game changer for embedding dynamic, responsive voice interactions in your applications, providing unprecedented realism, flexibility and latency. In fact, you're listening to one of their voices - right - now. If you are a developer looking to elevate user experience with natural voice interfaces, this is your solution. Visit up next dot fm slash eleven to check out their latest offerings. ...\n\nOpenAI is launching Daybreak, according to The Verge, a security initiative focused on detecting and patching vulnerabilities before attackers find them. The Verge says Daybreak uses OpenAI’s Codex Security AI agent, which launched in March, to build a threat model from an organization’s code and focus on likely attack paths.\n\nOpenAI also announced DeployCo, a new enterprise deployment company built to help organizations bring frontier AI into production and turn it into measurable business impact. That comes directly from OpenAI’s announcement, so for now the clean takeaway is the launch itself and the company’s framing around enterprise deployment.\n\nBackblaze said its network telemetry shows AI-driven neocloud traffic is changing infrastructure requirements, especially around what it calls massive high-throughput data flows between storage and compute. The immediate news hook is a conference presentation, but the broader point is that AI infrastructure pressure is increasingly about moving data fast enough to keep GPUs busy.\n\nThe Times of India reports Sarvam AI is arguing India cannot afford to remain just a consumer in the AI era and needs to build its own frontier-scale models. That’s one more sign that national AI strategy is increasingly being framed around sovereign model capacity, not just application layers.\n\nAnd one more: Forbes reports on Anthropic’s work on so-called natural language autoencoders, a new approach meant to help interpret what is happening inside generative AI systems and large language models. The big picture there is interpretability: getting closer to understanding how models represent concepts internally, instead of relying on whatever explanation a chatbot gives after the fact.\n\nBefore we wrap up, a quick note: this podcast is generated with the assistance of AI and is intended for informational purposes only. All referenced articles, research, and commentary remain the property of their original authors and publishers.\n\nAnd that's your briefing for today. Full source links are in the episode notes, and we'll be back tomorrow with what's up next!