UpNext AI

AI is moving into physical experimentation, scientific workflows, and high-stakes clinical documentation. This episode examines Outer Biosciences’ living-skin discovery platform, Inherent’s Faraday research agent, and a study showing why expert review remains vital for AI-generated anesthesia drafts.
Covered stories:
- Outer Biosciences uses living donated human skin and an AI feedback loop to identify potential skincare compounds.
- Inherent says its Faraday agent reproduced published scientific results better than larger Anthropic and OpenAI models in its evaluation.
- A 15-case feasibility study found clinically relevant errors in LLM-generated preoperative anesthesia drafts, reinforcing the need for expert correction.
- An anonymous model named Ox Alpha appears on OpenRouter.
- Reporting points to continued demand for lower-cost Anthropic models.
- The UAE and U.S. plan a military AI task force in Abu Dhabi.
- A look at the online “cursed AI” image phenomenon.
Sources:
- Outer Biosciences / TechCrunch: https://techcrunch.com/2026/08/21/michael-polansky-is-training-an-ai-model-on-skin-thats-still-alive/
- Inherent Faraday / TechCrunch: https://techcrunch.com/2026/08/22/inherent-founded-by-deepmind-alumni-says-its-ai-teammate-just-outperformed-anthropic-and-openai-at-replicating-research/
- Anesthesia workflow study: https://doi.org/10.1016/j.medcli.2026.107571
- Ox Alpha / TechCrunch: https://techcrunch.com/2026/08/23/whos-behind-the-new-stealth-model-ox-alpha/
- Anthropic model adoption discussion: https://simonwillison.net/2026/Aug/23/anthropics-best-ai-model-struggles-to-attract-users-as-cheaper-t/
- UAE-U.S. military AI task force / Gulf News: https://gulfnews.com/uae/uae-and-us-to-launch-worlds-first-bilateral-military-ai-task-force-1.500648985
- Cursed AI image gallery: https://www.boredpanda.com/cursed-ai-pictures/

What is UpNext AI?

Daily AI news and research, distilled. UpNext AI breaks down the most important developments in artificial intelligence—from major industry moves to cutting-edge papers.

Welcome to the UpNext AI podcast. It's Monday, August 24th, 2026, and here's what matters in AI today.

Our lead story takes AI development out of the data center and into the lab. TechCrunch reports that Michael Polansky’s startup, Outer Biosciences, has spent years developing a system to keep donated human skin tissue alive outside the body for more than a month. The company sources de-identified tissue from biobanks and brokers, with documented donor consent and institutional review board oversight, then maintains it with nutrients and waste removal.

The goal is not to treat patients. Outer is using living tissue as a more realistic testing environment for skincare compounds than short-lived samples, simple cell cultures, or animal models. Its AI model predicts which untested chemicals may help a particular skin function. The company tests those candidates in the living-tissue system, then feeds the results back into the model for the next round of predictions.

Polansky told TechCrunch that the system can observe processes that take weeks, including pigmentation changes, barrier repair, and responses after UVB damage. Outer is selling a faster discovery loop for cosmetic ingredients, rather than a finished consumer product. It has also begun research partnerships with beauty brands and a pharmaceutical partner studying severe skin rashes associated with certain cancer drugs.

The larger significance is the feedback loop. AI is often strongest when it can learn from abundant digital data. Outer’s approach tries to create a recurring source of biological evidence instead. Whether that produces commercially useful ingredients at scale remains to be seen, but it is a notable example of machine learning being paired with a purpose-built physical experimentation system.

That same push toward AI systems that do more than generate text is also behind our second story. London-based Inherent, founded by Google DeepMind alumni, has released an agent called Faraday. The company says Faraday can independently reproduce the findings of published scientific papers without being given the answer in advance.

Inherent says Faraday outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 on that task, while running on Qwen 3.6, a model with 27 billion parameters. The company has raised a 50 million dollar seed round and says its eventual aim is an AI scientist that can contribute to discovery across fields.

Replication is an important but narrower task than making new discoveries. Reproducing a paper asks whether an agent can select and run the right experiments to recover a known result; it does not establish that the system can identify an original, valuable research question. And the reported comparison is Inherent’s own claim, without independent benchmark details in the announcement.

Still, the design choice is worth watching. Inherent says it used reinforcement learning to reward useful experimental outcomes and what it calls research taste: choosing worthwhile experiments and designing them well. It also had Faraday use OpenAI’s GPT-5.5 Codex for coding, instead of building every tool itself. The takeaway for teams building agents is that model capability may increasingly depend on how well systems plan, use tools, and learn from experimental feedback—not just on the size of the underlying model.

Now, a research note on where that distinction becomes especially consequential: preoperative anesthesia assessment. A single-centre exploratory feasibility study examined an LLM-assisted workflow for drafting assessments for 15 complex anesthesia consultation cases. The researchers used DeepSeek-R1 to create structured drafts from de-identified case information, then had three experienced anesthesiologists review them for completeness, scientific plausibility, and attention to key risks.

The drafts were structurally complete, but the reviewers identified 32 clinically relevant errors across the 15 cases. Those errors included proposed anesthesia plans, American Society of Anesthesiologists physical-status classifications, and interpretations of abnormal tests. Senior anesthesiologists corrected the material before it was shown to first-year anesthesia residents.

That workflow matters because a well-formatted draft can create false confidence. The study found that, after expert correction, residents viewed the materials favorably, but their self-reported ability to spot the annotated errors remained limited. This was not a test of autonomous clinical care: there was no comparison group, blinding, or objective clinical outcome measure. Its crisp takeaway is that AI may be useful for assembling a preliminary clinical draft, but expert review is essential when the draft could influence high-stakes medical decisions.

...Are you building apps with voice? Elevate your app's voice capabilities with ElevenLabs. Their API is a game changer for embedding dynamic, responsive voice interactions in your applications, providing unprecedented realism, flexibility and latency. In fact, you're listening to one of their voices - right - now. If you are a developer looking to elevate user experience with natural voice interfaces, this is your solution. Visit up next dot fm slash eleven to check out their latest offerings. ...

A mysterious free model called Ox Alpha has appeared on OpenRouter, described as a reasoning model for coding, sustained agentic work, and production workloads. TechCrunch reports that its anonymous third-party provider has chosen to remain unidentified during the preview. The speculation around its origin is just that—speculation—so the meaningful development is the release of an anonymous model into a public testing channel.

On frontier-model economics, Simon Willison highlighted Financial Times reporting that Anthropic’s cheaper models are capturing more usage than its top-end offerings. Billing-data estimates from Ramp, which tracks spending across 70,000 companies using its credit cards, showed Anthropic model use spread across several releases rather than concentrated in the newest premium model. For buyers, price and task fit may matter more than always choosing the flagship.

The UAE and the United States are set to launch a bilateral military AI task force in Abu Dhabi. According to CENTCOM, the unit will apply AI to intelligence support, critical-infrastructure protection, and monitoring the regional security environment. It is another sign that AI deployment is becoming an operational focus of defense partnerships, not only a research priority.

And one lighter signal from the culture around generative AI: Bored Panda has collected 92 so-called cursed AI images shared by an online community. It is not a research finding, but it does reflect how visibly bad or uncanny outputs are becoming part of the public’s everyday experience of the technology.

Before we wrap up, a quick note: this podcast is generated with the assistance of AI and is intended for informational purposes only. All referenced articles, research, and commentary remain the property of their original authors and publishers.

If you enjoyed this episode, don't forget to subscribe, rate, and leave us a review! And that's your briefing for today. Full source links are in the episode notes, and we'll be back tomorrow with what's up next!