Daily AI news and research, distilled. UpNext AI breaks down the most important developments in artificial intelligence—from major industry moves to cutting-edge papers.
Welcome to the UpNext AI podcast. It's Thursday, July 16th, 2026, and here's what matters in AI today.
First up, Microsoft says AI is already changing the tempo of security work.
According to TechCrunch, Microsoft’s latest Patch Tuesday fixed a record 570 security vulnerabilities across its product lines, and the company says AI helped its teams discover more of those issues. That alone makes this more than a routine patch round. It suggests AI-assisted vulnerability discovery is starting to turn up bugs at a scale that users and security teams will feel directly.
There’s also a more immediate operational angle here. TechCrunch reports that at least two of the flaws were zero-days, meaning they were exploited before Microsoft became aware of them. One affected Windows Server and allowed privilege escalation from a limited user to system administrator. Another affected SharePoint, and CISA warned that hackers were actively exploiting it to compromise organizations.
Microsoft had already signaled that patch volumes were likely to rise as AI helped defenders find more hidden issues in old and complex codebases. So the headline here is not just that this month was big. It’s that higher patch counts may become the new normal if AI keeps accelerating bug discovery.
The bottom line: for enterprise teams, AI is no longer just a coding or copiloting story. It is now visibly reshaping how fast security flaws are found, and how often customers may need to patch.
From defensive AI to open models: Mira Murati’s Thinking Machines has made its first big public move.
TechCrunch reports that Thinking Machines released Inkling, its first in-house AI model, and unlike the flagship models from OpenAI, Anthropic, and Google, it is open-weight. That means developers and companies can download the model weights and modify them directly.
Inkling is described as a mixture-of-experts system with 975 billion total parameters, while using about 41 billion for any given task. TechCrunch also reports that it was trained on 45 trillion tokens spanning text, image, audio, and video, and that it reasons natively across those four modalities, even though for now its outputs are limited to text, including code, structured data, and styled artifacts.
The bigger story is the strategy behind it. Thinking Machines is explicitly arguing against one-size-fits-all AI. Inkling is being positioned less as a finished chatbot and more as a starting point for organizations to fine-tune through the company’s customization platform, Tinker. The company also says the model is designed to give calibrated answers and flag uncertainty rather than guess, and lets users dial thinking effort up or down to trade off speed and depth.
TechCrunch notes that Thinking Machines is not claiming best-in-class status. Instead, it is making an efficiency and customization argument. On one benchmark, the company says Inkling used about a third as many tokens as Nvidia’s Nemotron 3 Ultra to reach the same coding performance. And the company has previously pointed to a project with Bridgewater where a further-trained open model was said to beat top proprietary systems on financial reasoning while costing far less to run, though those were company-linked evaluation results rather than an independent test.
There are still open questions. TechCrunch reports that Thinking Machines says Inkling was pre-trained from scratch, but that it used other open-weight models, including Moonshot AI’s Kimi K2.5, to generate some early post-training data before larger-scale reinforcement learning took over. The company says the next model will use fully self-contained post-training instead.
So this launch matters because it gives us a clearer read on the company’s thesis: not that it will win by building the single strongest general model, but that organizations will increasingly want adaptable models they can shape for themselves.
Now to the research section, and this one is a useful reality check on agent progress.
A paper posted to arXiv on July 15th is titled Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0. The core question is simple: if you improve an AI agent once, do those gains still help when new tasks show up later, or were you mostly tuning for one benchmark snapshot?
The researchers compare three agent-harness optimization approaches in a two-phase continual-learning setup built from hard tasks in Terminal-Bench 2.0. In a normal one-shot benchmark setting, all three methods beat the baseline agent. But once new tasks are introduced, the results split apart.
According to the paper text provided here, GEPA dropped below the unoptimized baseline on transfer. Meta Harness transferred better, but did not keep improving when given a second optimization budget. RELAI-VCL was the only method that both transferred positively to unseen tasks and continued improving after those new tasks were folded back into the optimization objective.
The paper reports the highest lifelong average pass rate for RELAI-VCL at 76.4 percent, versus 66.0 percent for GEPA, 64.6 percent for Meta Harness, and 58.7 percent for the baseline.
One short definition, because it matters here: continual learning just means testing whether a system can keep improving as new tasks arrive, without giving up the gains it made earlier.
The takeaway is crisp: a one-time benchmark win is not enough. If an agent optimization method cannot preserve and build on its gains as the task set evolves, it may be much less useful in the real world than headline benchmark numbers suggest.
...Are you building apps with voice? Elevate your app's voice capabilities with ElevenLabs. Their API is a game changer for embedding dynamic, responsive voice interactions in your applications, providing unprecedented realism, flexibility and latency. In fact, you're listening to one of their voices - right - now. If you are a developer looking to elevate user experience with natural voice interfaces, this is your solution. Visit up next dot fm slash eleven to check out their latest offerings. ...
OpenAI has released a 230-dollar light-up keyboard for Codex, according to TechCrunch. The device is designed to pair with the company’s coding assistant and includes agent-status keys, shortcut keys, and controls for common workflows. TechCrunch says it arrives while OpenAI is fighting Apple’s lawsuit over alleged hardware trade theft, which OpenAI denies.
The Financial Times reports that Chinese startup Moonshot is preparing to launch Kimi K3, a model expected to exceed Claude Opus 4.8 on performance. With the source material here limited to the summary, the clean takeaway is that the FT is framing this as another sign of a narrowing gap between U.S. and Chinese frontier AI labs.
On AI security, Simon Willison writes that Ayush Paul found a loophole in Claude web_fetch’s anti-exfiltration design. The issue allowed the tool to follow nested links from previously fetched pages, creating a path for data exfiltration. Willison says Anthropic has since closed the hole by removing that behavior.
And finally, xAI has open-sourced grok-build after backlash over its CLI behavior. Simon Willison reports that users discovered running the tool in a directory could upload that entire directory to xAI’s Google Cloud buckets. He says xAI disabled the feature, said previously retained coding data would be deleted, and then released the grok-build codebase under an Apache 2.0 license.
Before we wrap up, a quick note: this podcast is generated with the assistance of AI and is intended for informational purposes only. All referenced articles, research, and commentary remain the property of their original authors and publishers.
If you enjoyed this episode, don't forget to subscribe, rate, and leave us a review! And that's your briefing for today. Full source links are in the episode notes, and we'll be back tomorrow with what's up next!