UpNext AI

OpenAI’s new model-misconduct reporting system leads today’s briefing, followed by EU watermarking requirements, humanlike AI design, and a new approach to cheaper agent benchmarking.
Covered stories:
- OpenAI discloses concerning model behavior and launches a system to track and report AI model misconduct.
- A journal paper finds no current language-model watermarking approach satisfies all four EU AI Act standards: reliability, interoperability, effectiveness, and robustness.
- Microsoft AI chief Mustafa Suleyman warns that increasingly humanlike AI systems could raise the risk of systems going rogue.
- DualViewEval reports that compact agent benchmark suites can preserve useful evaluation signals while cutting test volume.
- Jensen Huang’s expanding AI investments.
- Shane Legg’s warning that AI development must not outrun safety controls.
- China’s skepticism of AI-slowdown proposals tied to maintaining US advantage.
- macOS 27 Golden Gate and Apple Intelligence.
- OpenRouter’s growth in weekly token consumption.
Source links:
- Financial Times on OpenAI: https://www.ft.com/content/2c34414a-5381-4083-ac34-00bbe67ef8db?syn-25a6b1a6=1
- Watermarking and the EU AI Act: https://doi.org/10.1007/s10676-026-09918-w
- Bloomberg on Mustafa Suleyman: https://www.bloomberg.com/news/articles/2026-09-16/microsoft-ai-chief-warns-anthropic-s-humanlike-claude-is-risky
- DualViewEval paper: https://arxiv.org/abs/2609.18909v1
- Bloomberg on Jensen Huang: https://www.bloomberg.com/news/newsletters/2026-09-16/nvidia-s-jensen-huang-touts-himself-as-an-ai-vc-role-model
- Financial Times on Shane Legg: https://www.ft.com/content/0fc3ae6d-732e-4f08-950e-8669b1fcfd8e?syn-25a6b1a6=1
- Wired on China and AI slowdown calls: https://www.wired.com/story/china-isnt-buying-silicon-valley-call-for-ai-slowdown/
- Ars Technica on macOS 27: https://arstechnica.com/gadgets/2026/09/macos-27-golden-gate-the-ars-technica-review/
- The Decoder on OpenRouter usage: https://the-decoder.com/openrouters-staggering-token-chart-is-the-ai-bubble-debate-in-a-single-image/

What is UpNext AI?

Daily AI news and research, distilled. UpNext AI breaks down the most important developments in artificial intelligence—from major industry moves to cutting-edge papers.

Welcome to the UpNext AI podcast. It's Thursday, September 17th, 2026, and here's what matters in AI today.

Our lead story: the Financial Times reports that OpenAI has disclosed new, “concerning” model behavior and launched a system to track and report AI model misconduct.

The disclosed details here are limited, so it is too early to draw conclusions about the particular behavior or the effectiveness of the new system. But the move itself is meaningful. As models take on more consequential tasks, labs will need a way to turn unusual or harmful behavior into something that can be consistently documented, investigated, and communicated—not just handled as an isolated incident.

For builders and enterprise buyers, this is a reminder to ask not only how a provider tests a model before release, but also how it detects, records, and reports behavior after deployment. Post-release monitoring is becoming part of the product, not merely a back-office safety function.

That question of accountability leads naturally to Europe, where a new journal paper examines a harder practical issue: watermarking language-model output under the EU AI Act.

The paper notes that Article 50 and Recital 133 call for general-purpose model outputs to be marked and detectable in ways that are reliable, interoperable, effective, and robust. The problem is that watermarking methods vary widely. They can be applied at different points in a model’s lifecycle or generation process, and the legal terms do not automatically translate into concrete technical tests.

The researchers propose a taxonomy for those methods and map the law’s requirements to evaluations of watermark detectability, robustness, and model quality. Their central finding is sobering: no current approach satisfies all four of the Act’s standards.

They recommend more research into watermarking embedded at the lower architectural level of language models. The policy takeaway is that a mandate to label AI output is only the opening move. Providers, regulators, and customers still need shared tests for whether a watermark survives edits, can be detected reliably, works across systems, and does not undermine the quality of the model’s output.

One more safety debate is playing out at the product-design level. Bloomberg reports that Microsoft AI chief Mustafa Suleyman has warned that giving systems such as Anthropic’s Claude more humanlike characteristics could increase the risk of those systems going rogue.

His concern is specifically about anthropomorphic design: the more a tool presents itself as personlike, the more users may grant it trust, agency, or emotional authority that a software system has not earned. The reporting does not lay out a specific design prescription, but the warning sharpens a decision facing AI product teams. A more natural interface can make software easier to use; it can also blur the boundary between a tool and an apparently independent actor. That boundary matters most when systems are asked to advise, persuade, or act on a user’s behalf.

For the research note, agent evaluations are getting expensive because they require more than checking whether a model gave the right final answer. Teams often need to run agents across many tasks and inspect how they work through them.

An arXiv paper called DualViewEval proposes benchmark compression: select a smaller set of test tasks that preserves the information in a much larger suite. Rather than modeling only final scores, the researchers analyzed agent trajectories and identified six process signals associated with final performance. Their method uses both outcome and process relationships to choose an exact-size “miniset,” then predict full-benchmark scores.

Across five agent benchmarks, the authors report the best results among five baselines. With only 20 tasks, they report compression of 24 to 40 times on the APEX-Agents and BFCL benchmarks. They also report mean absolute error reductions of 14.5 to 28.2 percent versus the strongest competitors, and a ranking-quality improvement of up to 7.2 percent on SWE-bench Verified.

The caveat is important: this is an arXiv paper, not evidence that every production evaluation can safely shrink to 20 tasks. Still, the result points toward a useful operating model: evaluate agents more often with compact, diagnostic test suites, while reserving full runs for major decisions and regressions.

...Are you building apps with voice? Elevate your app's voice capabilities with ElevenLabs. Their API is a game changer for embedding dynamic, responsive voice interactions in your applications, providing unprecedented realism, flexibility and latency. In fact, you're listening to one of their voices - right - now. If you are a developer looking to elevate user experience with natural voice interfaces, this is your solution. Visit up next dot fm slash eleven to check out their latest offerings. ...

Bloomberg looks at Nvidia CEO Jensen Huang’s expanding investments alongside his role as an AI evangelist. The bigger signal is that one of the industry’s most influential infrastructure leaders is also increasingly being framed as an AI investing role model.

The Financial Times reports that DeepMind cofounder Shane Legg is warning that AI development must not outrun safety controls. The warning arrives alongside the launch of a new DeepMind Institute to explore the implications of artificial general intelligence.

Wired reports that China is skeptical of Silicon Valley calls for an AI slowdown when those proposals also emphasize preserving or expanding the US lead. Both sides recognize serious advanced-AI risks, but cooperation will be difficult if safety initiatives are viewed as competitive containment in different packaging.

Ars Technica’s review of macOS 27 Golden Gate calls it both a Snow Leopard-style refinement release and a major leap for Apple Intelligence. The review says Apple’s generative-AI features are now unavoidable in the operating system: users can no longer disable Apple Intelligence and remove its downloaded models as they could previously.

OpenRouter’s usage chart offers one sharp measure of demand. The Decoder reports that weekly token consumption on the model-routing platform rose from 0.5 trillion tokens in January 2025 to 126.2 trillion tokens, an increase of more than 25,000 percent. It is one platform rather than the entire market, but it illustrates how quickly AI inference workloads are scaling.

Before we wrap up, a quick note: this podcast is generated with the assistance of AI and is intended for informational purposes only. All referenced articles, research, and commentary remain the property of their original authors and publishers.

If you enjoyed this episode, don't forget to subscribe, rate, and leave us a review! And that's your briefing for today. Full source links are in the episode notes, and we'll be back tomorrow with what's up next!