Daily AI news and research, distilled. UpNext AI breaks down the most important developments in artificial intelligence—from major industry moves to cutting-edge papers.
Welcome to the UpNext AI podcast. It's Monday, May 4th, 2026, and here's what matters in AI today.\n\nOur top story: TechCrunch reports on a new study from a team led by physicians and computer scientists at Harvard Medical School and Beth Israel Deaconess Medical Center, published in Science, looking at how large language models perform in medical settings, including real emergency-room cases.\n\nThe headline result is striking. In one experiment, the researchers looked at 76 patients who came through the Beth Israel emergency room and compared diagnoses from two internal medicine attending physicians with diagnoses generated by OpenAI’s o1 and 4o models. Those diagnoses were then evaluated by two other attending physicians who were blinded to whether the answers came from humans or AI.\n\nAccording to the study, o1 performed either slightly better than or on par with the physicians and 4o at each diagnostic touchpoint, with the biggest difference at the initial ER triage stage, when there is the least information and the most urgency. TechCrunch says Harvard’s press release emphasized that the models were given the same electronic medical record information available at the time, without preprocessing. In triage, the o1 model produced the exact or very close diagnosis in 67 percent of cases, compared with 55 percent for one physician and 50 percent for the other.\n\nThat does not mean AI is ready to run the emergency room. The study itself called for prospective real-world trials, and TechCrunch notes one important critique from emergency physician Kristen Panthagani: the study compared the model with internal medicine physicians, not ER physicians, and in emergency care the first job is often to identify immediately dangerous conditions, not simply guess the final diagnosis. Still, this is the kind of result that matters because it pushes the conversation past demos and into a real high-stakes workflow. The question is no longer whether AI can produce plausible medical language. It is whether, in carefully bounded settings, it can become a reliable second set of eyes.\n\nFor the second act today, we’re looking at what the AI builder ecosystem seems to think matters next. Latent Space’s AINews is out with a Wave 2 call for speakers for this summer’s AI Engineer World’s Fair, and while this is an event announcement, it doubles as a useful snapshot of where the conversation is moving.\n\nThe organizers say this year’s event is adding an entire day of talks and are specifically soliciting projects in autoresearch, memory, world models, what they call tasteful tokenmaxxing, agentic commerce, and vertical AI in law, healthcare, go-to-market, and finance. They also call out robotics, including free expo floor space for strong robotics demos, and a new startup battlefield for pre-Series A companies.\n\nRead one way, this is just programming. Read another way, it’s a market signal. The emphasis is not only on better base models; it’s on systems that remember, act, transact, and fit into domain-specific workflows. Even the terminology tells you something. Memory and world models point to more persistent and grounded agents. Agentic commerce suggests infrastructure for software that can actually buy data, call APIs, and coordinate with other systems. And vertical AI keeps surfacing because generic capability is only part of the story; the real test is whether AI can map onto the constraints of healthcare, finance, law, and sales.\n\nSo the broader takeaway here is that the center of gravity keeps moving away from chatbot novelty and toward applied engineering. The frontier is increasingly about runtime design, domain fit, and whether these systems can operate inside real businesses without creating more chaos than value.\n\nNow to today’s research pick. A new Scientific Reports paper audits five AI chatbots on concussion health advice, comparing retrieval-augmented systems with standard pretrained models.\n\nIn plain English, this is a very practical test. The researchers were not asking whether a chatbot can sound informative. They were asking whether it can give reliable, understandable advice to someone looking up a common medical issue. They used 11 high-volume patient questions from Google Trends, ran them through five platforms in a zero-shot setup, and had two blinded neurosurgeons score the answers against the 2023 Amsterdam Consensus Statement.\n\nThe paper used several quality measures. DISCERN and EQIP looked at treatment and information quality, GQS measured overall content quality, and JAMA benchmarks looked at transparency. Readability was also tested with standard reading-level metrics.\n\nThe main result: reliability differed significantly across systems, and the retrieval-augmented model Perplexity Pro scored highest on the DISCERN and EQIP measures, outperforming foundational models including ChatGPT and Gemini in this audit. The authors suggest that gap is likely related to retrieval augmentation. But two problems remained. Transparency was uniformly low, and readability was still too difficult. All of the models exceeded a sixth-grade reading level, and most were above tenth-grade. Perplexity Pro was the easiest to read in the group, but still not really plain-language by default.\n\nAnd that is the useful lesson here. Better retrieval can improve the factual quality of health answers, but it does not automatically make those answers transparent or easy for patients to understand. Bottom line: if a chatbot is going anywhere near patient education, fluency is not enough — you need reliability checks, readability checks, and human review.\n\n...Are you building apps with voice? Elevate your app's voice capabilities with ElevenLabs. Their API is a game changer for embedding dynamic, responsive voice interactions in your applications, providing unprecedented realism, flexibility and latency. In fact, you're listening to one of their voices - right - now. If you are a developer looking to elevate user experience with natural voice interfaces, this is your solution. Visit up next dot fm slash eleven to check out their latest offerings. ...\n\nFirst headline: CBC reports that families of victims of the Tumbler Ridge, British Columbia school shooting who are suing OpenAI may face major legal hurdles. The core issue, according to the report, is whether OpenAI had a duty to warn authorities about the shooter’s interactions with ChatGPT, and whether a failure to act can be shown to have caused the attack. CBC says legal experts see the case as unusually difficult because it turns on questions like special duty, causation, and whether Section 230 protections might apply.\n\nAnd finally, a lighter developer note. Simon Willison shared a small tool called iNaturalist Sightings that he says he built entirely on his phone using Claude Code for web. The app groups observations from two iNaturalist accounts by time and place, using a Python CLI, a Git scraping repository, and a simple front end that fetches JSON from GitHub. It’s a nice example of what AI-assisted development looks like when it’s pragmatic, personal, and fast rather than grandiose.\n\nBefore we wrap up, a quick note: this podcast is generated with the assistance of AI and is intended for informational purposes only. All referenced articles, research, and commentary remain the property of their original authors and publishers.\n\nAnd that's your briefing for today. Full source links are in the episode notes, and we'll be back tomorrow with what's up next!