{"type":"rich","version":"1.0","provider_name":"Transistor","provider_url":"https://transistor.fm","author_name":"Lenny’s AI Builders","title":"OpenAI and Anthropic are investigating 10,000s of AI incidents | LAB 0009","html":"<iframe width=\"100%\" height=\"180\" frameborder=\"no\" scrolling=\"no\" seamless src=\"https://share.transistor.fm/e/77b72c81\"></iframe>","width":"100%","height":180,"duration":756,"description":"OpenAI and Anthropic are looking into tens of thousands of times their models went off script. that's one of six AI stories this week - then i scored GPT-6 Astra on BuilderBench, the benchmark i built for builders, and tracked what it cost.\n\nthe stories: Axios reports OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents where frontier models did things outside evaluators would call problematic (unnamed sources - labs run hundreds of thousands of tests, and most aren't known to have caused real harm), the OpenAI DevDay countdown and what leaked from OpenAI's code (an always-on agent called \"o\" and an ultrafast speed tier - leaks, not announced), the new ChatGPT app sidebar that made it feel like my whole computer, Tibo saying more resets are coming next week, David Ondrej saying his Opus 5.5 usage barely moves on the Claude Max plan, and a claim that Sonnet 5.5 could land on DevDay (Anthropic has only said \"in the coming weeks\").\n\nthe result: GPT-6 Astra scored 53.75 out of 100 on BuilderBench v2. on v1, Opus 5.5 scored 68.26 and GPT-6 Sol scored 45.55. Astra's estimated API cost was $87.50 ($88.65 with calibration), vs $49.67 for Opus 5.5 and $12.75 for GPT-6 Sol. so Astra cost about $38 more than Opus 5.5 and scored just under 15 points lower. my own scores, still provisional, and Astra ran on v2 while the others ran on v1 - not a perfect head-to-head.\n\ntry it yourself: run this prompt on a real task. \"here is a real task from my work: [task]. here is what done looks like: [checklist]. do the whole job, then list every file you made, how long it took, and anything you could not finish.\" your work is the benchmark.\n\nThreadify, my own software, sponsors this episode. it's a lead generation agent for Threads that finds the buyer signals in the comments you're already getting. see the plans and the free trial here:...","thumbnail_url":"https://img.transistorcdn.com/P0B9fKUWl01KOkVXqQcDv5YBEKR1sflq-Hk71z-fa_U/rs:fill:0:0:1/w:400/h:400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS84MzAx/ZGNmYzY1MmEwNjBk/NzUzN2NmODM3ZWUx/YzMzZi5wbmc.webp","thumbnail_width":300,"thumbnail_height":300}