Daily AI news and research, distilled. UpNext AI breaks down the most important developments in artificial intelligence—from major industry moves to cutting-edge papers.
Welcome to the UpNext AI podcast. It's Friday, July 17th, 2026, and here's what matters in AI today.
We start with Kimi K3. Simon Willison reports that Chinese AI lab Moonshot AI has announced Kimi K3, which the company describes as its most capable model to date, with 2.8 trillion parameters. It’s available now through Moonshot’s website and API, and an open-weight release is promised by July 27. Moonshot is calling it the first open 3T-class model, and according to Willison’s write-up, its self-reported benchmarks put K3 mostly ahead of Claude Opus 4.8 max and GPT-5.5 high, though still behind Claude Fable 5 and GPT-5.6 Sol. Willison also points to third-party notes from Artificial Analysis, which said K3 reached an overall Elo of 1547 on a private long-horizon knowledge-work evaluation, up 732 points from Kimi K2.6 and behind only Claude Fable 5. Artificial Analysis also said cost per task was about 94 cents, close to GPT-5.6 Sol at $1.04, while K3 used 21 percent fewer output tokens than K2.6 on its intelligence index. Willison says K3 is priced at $3 per million input tokens and $15 per million output tokens, putting it around Anthropic’s Claude Sonnet pricing band and making it, in his description, the most expensive model released by a Chinese AI lab to date. The bigger signal here is that the open-model race is no longer just about being cheaper. Moonshot is making a serious play on scale, coding credibility, and eventual open-weight access, while suggesting frontier-class performance may come with frontier-class pricing.
Our second story is about AI inside a very traditional business category: travel. TechCrunch reports that Fora, an AI-powered travel agency, has raised a $60 million Series D led by Forerunner and Tactile Ventures, valuing the company at $1 billion. The company says it has now raised $138.5 million in total. Fora, founded in 2021, helps people become travel agents with tools for client communication and trip planning, while also helping travelers find and work with advisors. TechCrunch says part of the new funding will go toward expanding Fora’s AI assistant, Via, which helps agents with administrative work like research and itinerary building. The interesting angle here is that Fora is not pitching AI as a replacement for human advisors. It’s pitching AI as leverage for them. TechCrunch also reports that agents on the platform have booked more than $3 billion in travel since launch, with most agents new to travel advising. So this looks like another case where AI is being used to widen access to skilled service work and improve productivity, not just automate it away.
For the research note, earlier this week a new arXiv paper asked: can we trust item response theory for AI evaluation? Item response theory, or IRT, is a statistical method used to analyze individual test questions so researchers can estimate model capability, rank systems, pick informative examples, and judge benchmark quality. The paper argues that AI benchmark data often looks very different from the human-testing setups where standard IRT tools were developed. The authors say AI benchmarks tend to have fewer evaluated models, many more items, and capability distributions that can be skewed, clustered, or multimodal. Using data derived from six widely used LLM benchmarks, they simulated response matrices under three common IRT models and compared four estimation approaches across 18,000 simulation conditions. Their conclusion is a caution flag: classical estimators can become infeasible at large benchmark scales, while more scalable estimators can produce unreliable item-level and ranking inferences when the model set is small or not normally distributed. Bottom line: if a benchmark uses IRT, that method may be useful, but it is not automatically trustworthy without the right sample sizes and diagnostics.
...Are you building apps with voice? Elevate your app's voice capabilities with ElevenLabs. Their API is a game changer for embedding dynamic, responsive voice interactions in your applications, providing unprecedented realism, flexibility and latency. In fact, you're listening to one of their voices - right - now. If you are a developer looking to elevate user experience with natural voice interfaces, this is your solution. Visit up next dot fm slash eleven to check out their latest offerings. ...
Google is renaming NotebookLM to Gemini Notebook. TechCrunch reports the company is also adding code execution for data analysis inside each notebook’s secure container, and says users will soon be able to access notebooks through AI Mode in Search.
Thinking Machines Lab has released its first open-weights model, Inkling. Simon Willison describes it as an Apache-2.0 licensed multimodal mixture-of-experts model with 975 billion total parameters and 41 billion active parameters, trained on 45 trillion tokens.
Netflix says generative AI is now showing up across a meaningful chunk of its catalog. The Verge reports that around 300 titles on the platform used generative AI, with most of that use happening in post-production, according to Netflix’s second-quarter earnings report.
And in Europe, Ars Technica reports that the EU will force Google to share search data and open up AI on Android under legally binding Digital Markets Act measures. Google says those changes could endanger privacy and security.
Before we wrap up, a quick note: this podcast is generated with the assistance of AI and is intended for informational purposes only. All referenced articles, research, and commentary remain the property of their original authors and publishers.
If you enjoyed this episode, don't forget to subscribe, rate, and leave us a review! And that's your briefing for today. Full source links are in the episode notes, and we'll be back Monday with what's up next!