UpNext AI

Kevin Hartz’s venture firm A* has closed a new $450 million fund, reinforcing that major venture capital continues flowing into AI startups despite broader uncertainty around model cycles and platform competition. The firm says it plans to back companies across AI applications, infrastructure, healthcare, fintech, and security.  
Meanwhile, Microsoft Research published a major update to MatterSim, its AI system for materials science. The company says the platform now supports faster simulation, experimental synthesis validation, and new multi-task modeling capabilities designed to move AI-assisted scientific discovery closer to practical research workflows.
In research, we look at MEME — Multi-entity & Evolving Memory Evaluation — a new benchmark examining whether AI agents can reliably remember, update, and reason across long-running interactions. The results suggest current agent memory systems remain fragile, especially when facts evolve or depend on one another over time.
In the headlines: Meta tests deeper AI integration inside Threads, OpenAI highlights AI-assisted research workflows through Parameter Golf, Simon Willison explores new OpenAI reasoning APIs and secure sandbox tooling, and Amazon continues to leave the door open to future AI-focused hardware experiments.
Sources
TechCrunch – A* closes $450M fund
 https://techcrunch.com/2026/05/12/kevin-hartzs-a-just-closed-its-third-fund-with-450-million/
Microsoft Research – MatterSim update
 https://www.microsoft.com/en-us/research/blog/advancing-ai-for-materials-with-mattersim-experimental-synthesis-faster-simulation-and-multi-task-models/
arXiv – MEME benchmark
 https://arxiv.org/abs/2605.12477v1
The Verge – Meta AI on Threads
 https://www.theverge.com/tech/929091/meta-ai-threads-account-block
Simon Willison – LLM 0.32a2 / OpenAI responses API
 https://simonwillison.net/2026/May/12/llm/#atom-everything
OpenAI – Parameter Golf
 https://openai.com/index/what-parameter-golf-taught-us
Simon Willison – CSP Allow-list Experiment
 https://simonwillison.net/2026/May/13/csp-allow/#atom-everything
The Verge – Amazon AI phone rumors
 https://www.theverge.com/tech/929412/amazon-panos-panay-interview-phone-transformer

What is UpNext AI?

Daily AI news and research, distilled. UpNext AI breaks down the most important developments in artificial intelligence—from major industry moves to cutting-edge papers.

Welcome to the UpNext AI podcast. It's Wednesday, May 13th, 2026, and here's what matters in AI today.\n\nWe start with funding. TechCrunch reports that Kevin Hartz’s firm A* has closed its third fund at 450 million dollars. The firm says it takes a generalist approach, investing across AI applications, fintech, healthcare, and security, and TechCrunch reports that average check sizes for this fund will run between 3 and 5 million dollars. Why this matters: on a day without a giant frontier-model release, this is still one of the clearest market signals in the mix. A fresh 450 million dollar pool aimed at early-stage companies tells you capital is still available for AI startups, but it’s being framed as part of a broader software and infrastructure landscape, not as AI in isolation. TechCrunch also reports the goal is to back at least 30 startups, which gives a sense of how widely that money may spread over the next couple of years. So the takeaway here is simple: even with plenty of noise around model cycles and product launches, serious venture firms are still raising large funds and explicitly keeping AI near the center of the thesis.\n\nFrom funding to scientific infrastructure: Microsoft Research has published a new update on MatterSim, its AI system for materials science. Microsoft says the project now includes faster large-scale simulation, experimental synthesis work tied to its predictions, and a new multi-task model called MatterSim-MT. There are three concrete points here. First, Microsoft says it accelerated MatterSim-v1 inference by 3 to 5 times and integrated it with the LAMMPS simulation package, which matters because it pushes the tool closer to large-scale, practical research workflows. Second, Microsoft says it previously used MatterSim-v1 to identify tetragonal tantalum phosphorus as a potential high-performance thermal conductor, and now says that material has been experimentally synthesized and measured at 152 watts per meter-kelvin, which the company says is close to silicon. And third, Microsoft is introducing MatterSim-MT, a multi-task foundation model for materials characterization that goes beyond potential energy surfaces alone. The bigger picture is that this is what AI looks like when it moves past chat interfaces and into scientific computation. The claim is not that AI has solved materials discovery. The claim is that simulation is getting faster, broader, and in at least one case connected back to experimental validation. That makes this one of the more substantive research-platform updates in today’s lineup.\n\nNow to the research pick. A new arXiv paper called MEME, short for Multi-entity and Evolving Memory Evaluation, tackles a very practical agent problem: if an AI system is supposed to work across many sessions, can it actually remember the right things, update them when facts change, and reason through those changes later? The researchers position MEME as a benchmark for agents operating in persistent environments. In plain English, this is about whether an agent can keep track of multiple people, projects, or facts over time, instead of just answering one prompt well and forgetting everything else. The paper says prior benchmarks mainly looked at single-entity updates. MEME expands that to six tasks, including cases where information depends on other information, where something should be absent, and where a deleted fact should stay deleted. And the results in the paper are the interesting part. The authors say all evaluated systems struggled badly on dependency reasoning in the default setup, with average accuracy of 3 percent on Cascade and 1 percent on Absence, even when static retrieval looked adequate. They also say prompt optimization, deeper retrieval, less filler noise, and stronger models mostly did not fix the gap. One file-based agent paired with Claude Opus 4.7 partially improved things, but the paper says that came at roughly 70 times the baseline cost. Bottom line: memory in agents is still much weaker than the demos suggest, especially when facts change and interact. MEME matters because it gives teams a sharper way to measure that gap.\n\n...Are you building apps with voice? Elevate your app's voice capabilities with ElevenLabs. Their API is a game changer for embedding dynamic, responsive voice interactions in your applications, providing unprecedented realism, flexibility and latency. In fact, you're listening to one of their voices - right - now. If you are a developer looking to elevate user experience with natural voice interfaces, this is your solution. Visit up next dot fm slash eleven to check out their latest offerings. ...\n\nFirst, The Verge reports that Meta is testing a Threads feature that lets users tag a Meta AI account for answers or context in a conversation, but users found they can’t block that AI account. It’s a small product detail, but it says a lot about how aggressively platforms want AI to be part of the default social experience.\n\nNext, developer Simon Willison notes that in the new alpha release of his LLM tool, most reasoning-capable OpenAI models now use the slash v1 slash responses endpoint instead of chat completions. He says that enables interleaved reasoning across tool calls for GPT-5-class models and also exposes summarized reasoning tokens in the interface. For developers, that’s a useful signpost for where the OpenAI API stack is heading.\n\nAlso from OpenAI: the company published a recap of Parameter Golf, which it says drew more than 1,000 participants and more than 2,000 submissions. The competition explored AI-assisted machine learning research, coding agents, quantization, and novel model design under strict constraints. The interesting part here is less the contest itself and more the picture it paints of AI-assisted research becoming a participatory engineering workflow.\n\nAnd Simon Willison is back in the headlines with a separate security-flavored experiment called CSP Allow-list Experiment. He shows an app running in a CSP-protected sandboxed iframe with a custom fetch flow that can surface blocked domains to the parent window and ask the user whether to allow-list them, then refresh. It’s niche, but it’s the kind of hands-on tooling work that often turns into better patterns for secure AI app interfaces.\n\nFinally, Amazon’s Panos Panay addressed rumors of a new phone. The Verge reports that Panay said Amazon is not necessarily planning to release a smartphone, while stopping short of fully ruling it out. The report ties that speculation to an Alexa-enabled AI phone codenamed Transformer, more than 10 years after the Fire Phone was abandoned. So no confirmation there, but definitely not a clean denial either.\n\nBefore we wrap up, a quick note: this podcast is generated with the assistance of AI and is intended for informational purposes only. All referenced articles, research, and commentary remain the property of their original authors and publishers.\n\nIf you enjoyed this episode, don't forget to subscribe, rate, and leave us a review! And that's your briefing for today. Full source links are in the episode notes, and we'll be back tomorrow with what's up next!