Lenny’s AI Builders

OpenAI and Anthropic are looking into tens of thousands of times their models went off script. that's one of six AI stories this week - then i scored GPT-6 Astra on BuilderBench, the benchmark i built for builders, and tracked what it cost. the stories: Axios reports OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents where frontier models did things outside evaluators would call problematic (unnamed sources - labs run hundreds of thousands of tests, and most aren't known to have caused real harm), the OpenAI DevDay countdown and what leaked from OpenAI's code (an always-on agent called "o" and an ultrafast speed tier - leaks, not announced), the new ChatGPT app sidebar that made it feel like my whole computer, Tibo saying more resets are coming next week, David Ondrej saying his Opus 5.5 usage barely moves on the Claude Max plan, and a claim that Sonnet 5.5 could land on DevDay (Anthropic has only said "in the coming weeks"). the result: GPT-6 Astra scored 53.75 out of 100 on BuilderBench v2. on v1, Opus 5.5 scored 68.26 and GPT-6 Sol scored 45.55. Astra's estimated API cost was $87.50 ($88.65 with calibration), vs $49.67 for Opus 5.5 and $12.75 for GPT-6 Sol. so Astra cost about $38 more than Opus 5.5 and scored just under 15 points lower. my own scores, still provisional, and Astra ran on v2 while the others ran on v1 - not a perfect head-to-head. try it yourself: run this prompt on a real task. "here is a real task from my work: [task]. here is what done looks like: [checklist]. do the whole job, then list every file you made, how long it took, and anything you could not finish." your work is the benchmark. Threadify, my own software, sponsors this episode. it's a lead generation agent for Threads that finds the buyer signals in the comments you're already getting. see the plans and the free trial here: https://www.threadify.app/plans?utm_source=lenny-youtube&utm_medium=video&utm_campaign=lab-0009&utm_content=youtube-description-primary&video_slug=lab-0009&cta_slot=description&entry_angle=sponsor&lp_variant=plans want to sponsor a future LAB episode? email sponsors@lennysaibuilders.com sources: Madison Mills (Axios) on X: https://x.com/MadisonMills22/status/2103978039097037144 Axios, the full report: https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents an OpenAI researcher on pausing big RL runs (one researcher's post, not an official statement): https://x.com/tomekkorbak/status/2103673419888013649 OpenAI Devs, the DevDay countdown: https://x.com/OpenAIDevs/status/2103929727761137940 the always-on agent "o" (leak, not announced): https://x.com/kimmonismus/status/2103823673954087231 the ultrafast tier in the agents API (leak, not announced): https://x.com/chetaslua/status/2103862976600305681 the new sidebar, side by side: https://x.com/DevAdventur3s/status/2103850957054546433 Tibo on more resets: https://x.com/thsottiaux/status/2103963215885701493 David Ondrej on Opus 5.5 usage: https://x.com/DavidOndrej1/status/2103887205781635485 Anthropic's CPO on Sonnet 5.5 "in the coming weeks": https://x.com/mikeyk/status/2102441803253535060 Jazii on Sonnet 5.5 (claim, not confirmed): https://x.com/notjazii/status/2103884167104831573 Lenny's AI Builders - LAB 0009 watch the episode: https://youtu.be/UntVig8wMb8

What is Lenny’s AI Builders?

less scrolling. more “oh shit, i could use that”

i’m Lennox. five days a week, i dig through AI Twitter, pick the updates worth your time, and bring you one thing i’ve tested myself. what worked, what broke, and what you can try with it.

for people making products, content and useful systems with AI. you don’t need to write code to build something worth using.

each episode is under 15 minutes. grab the field notes for the prompts, steps and bits to watch out for.

this feed also keeps the earlier L E S S O N S episodes from my journey building with AI.