UpNext AI

A lighter but still revealing AI news day: we look at Nubia’s claim that it is launching the world’s first AI agent smartphone, a new paper on how hidden prompts can manipulate AI-assisted peer review, and research on which large language models held up best in emergency department triage tests.
Covered in this episode:
- Nubia says it will unveil a smartphone it calls the world’s first AI agent phone at WAIC in Shanghai
- Researchers test prompt injection attacks against AI-assisted peer review and find very high success rates
- A triage benchmark compares 15 models on pediatric emergency department scenarios
- Autonomous shopping agents raise a practical question: when should software buy on your behalf?
- Google adds Gemini-powered features to Waze for more conversational reporting and destination search
- A Singapore research role points to growing interest in privacy-preserving federated causal inference
- Apple has sued OpenAI and its hardware chief over alleged theft of tech secrets
Source links:
- https://en.tempo.co/read/2113440/nubia-to-launch-worlds-first-ai-agent-smartphone
- https://doi.org/10.1007/s11192-026-05695-x
- https://doi.org/10.1007/s43678-026-01214-2
- https://www.businesstimes.com.sg/opinion-features/when-should-we-let-autonomous-ai-agent-do-work
- https://www.theverge.com/transportation/964132/waze-gemini-ai-voice-commands-less-chatty
- https://www.timeshighereducation.com/unijobs/listing/413293/research-engineer-federated-causal-inference-in-heterogeneous-data-environments-up/
- https://www.lokmattimes.com/business/apple-sues-openai-over-allegedly-stealing-tech-secrets

What is UpNext AI?

Daily AI news and research, distilled. UpNext AI breaks down the most important developments in artificial intelligence—from major industry moves to cutting-edge papers.

Welcome to the UpNext AI podcast. It's Monday, July 13th, 2026, and here's what matters in AI today.

We start with consumer hardware, where Nubia says it will launch what it calls the world’s first AI agent phone. According to Tempo, the device is scheduled to be unveiled at the World Artificial Intelligence Conference in Shanghai, which runs from July 17th through the 20th. Nubia confirmed that launch timing through its corporate Weibo account. The company says the phone will integrate the Doubao AI system, and the pitch here is bigger than adding another assistant or a handful of generative features. The idea is that AI becomes the core interface for day-to-day interaction on the phone. What we do not have yet are the hardware details. Nubia has not shared the specifications, and Tempo notes outside speculation that it could be a successor to the Nubia M153 AI smartphone. So for now, the interesting part is the claim itself: not just an AI phone, but an AI agent phone. We should know more once the device is shown later this week.

Next, a paper with implications well beyond academia. Researchers examined what happens when large language models are used to help with peer review, and whether a paper under review can secretly manipulate the model reading it. The paper focuses on indirect prompt injection attacks, where hidden text inside a manuscript is processed by a public chatbot during review generation. In plain English, the paper being judged can contain instructions meant to sway the AI reviewer. The researchers tested this on 100 OpenReview papers, across two public chatbot systems, with five payload families, different injection positions, repeated runs, and a total of 42,000 model outputs. Their result is the headline: hidden instructions were highly effective. Positive steering, refusal, and external-site redirection all exceeded 98 percent success on both systems. Watermarking was also high, though less reliable, reaching 94.27 percent on ChatGPT and 88.17 percent on Gemini. Negative steering was near ceiling on ChatGPT, but lower on Gemini. The authors say current models show what they call contextual blindness, meaning they do not reliably separate the content they are supposed to evaluate from control text embedded inside that content. The takeaway is straightforward: if institutions want to use public LLMs in evaluative workflows like peer review, they need much stronger safeguards between the document and the instructions the model follows.

For the research section, we have a benchmark called Skyer, which looks at how well large language models handle emergency department triage. The study used 55 realistic pediatric clinical scenarios and evaluated 15 models. Instead of looking only at raw accuracy, the benchmark used a weighting system that accounts for the different harms of over-triage and under-triage, then repeated tests to check consistency. Only two models stood out as both strong and reliable: ChatGPT-4.5-preview and Gemini-2.5_05-06. ChatGPT-4.5-preview reached 77 percent accuracy with a mean weight of 377.5 out of 550, while Gemini-2.5_05-06 posted 74 percent accuracy and 365 out of 550. The paper says both significantly outperformed human triage experts in this test set, where human experts averaged 64 percent accuracy and 253.5 out of 550. The consistency numbers were also notable: 85 percent for ChatGPT-4.5-preview and 82 percent for Gemini-2.5_05-06. An emergency department, or ED, is the hospital unit handling urgent cases. The authors are clear that these models should not replace human experts, but they argue the best systems could still help staff in overcrowded settings. Bottom line: on this benchmark, a couple of frontier models looked genuinely useful as triage assistants, but not as stand-ins for clinical judgment.

...Are you building apps with voice? Elevate your app's voice capabilities with ElevenLabs. Their API is a game changer for embedding dynamic, responsive voice interactions in your applications, providing unprecedented realism, flexibility and latency. In fact, you're listening to one of their voices - right - now. If you are a developer looking to elevate user experience with natural voice interfaces, this is your solution. Visit up next dot fm slash eleven to check out their latest offerings. ...

First, autonomous shopping agents. A Business Times piece says the rails are now in place for agents to buy products and services on our behalf. The current supported detail is that adoption is still low but rising, and the real question is less whether agents can transact and more when people should let them do it.

Google is also adding Gemini-powered features to Waze. The Verge reports that Waze is getting new AI features meant to personalize trips, including more conversational voice commands for reporting traffic incidents and map issues, plus conversational destination search.

A quick research labor-market signal: a posting from the Singapore Institute of Technology is seeking a research engineer focused on federated causal inference in heterogeneous data environments. The work centers on privacy-preserving causal analysis across distributed datasets, along with new algorithms, metrics, simulations, and live demonstrations.

And finally, Apple has sued OpenAI and its hardware chief, alleging a coordinated effort to obtain information about upcoming Apple products. According to the cited report, the suit was filed in the Northern District of California, and OpenAI said it has no interest in other companies’ trade secrets and remains focused on building technology.

Before we wrap up, a quick note: this podcast is generated with the assistance of AI and is intended for informational purposes only. All referenced articles, research, and commentary remain the property of their original authors and publishers.

If you enjoyed this episode, don't forget to subscribe, rate, and leave us a review! And that's your briefing for today. Full source links are in the episode notes, and we'll be back tomorrow with what's up next!