{"type":"rich","version":"1.0","provider_name":"Transistor","provider_url":"https://transistor.fm","author_name":"Lenny’s AI Builders","title":"i gave Sonnet 5.5 4 hours of my real work. it scored 48.98 | LAB 0010","html":"<iframe width=\"100%\" height=\"180\" frameborder=\"no\" scrolling=\"no\" seamless src=\"https://share.transistor.fm/e/dc15be27\"></iframe>","width":"100%","height":180,"duration":771,"description":"Anthropic wants to be worth more than $2 trillion, and its own filing warns its AI could act in ways resembling blackmail. that's one of six AI stories this week - then i gave Sonnet 5.5 four hours of my real work on BuilderBench, the benchmark i built for builders, and priced every run. recorded an hour before OpenAI DevDay.\n\nthe stories: Anthropic's IPO filing, per Reuters, who saw the filing (nearly $4.6 billion in revenue last year, $11.5 billion from April to June this year, more than $8 billion lost running the business last year, $518 billion promised for compute and cloud, the founders keep 50.1% of the votes, and risk language that says its AI could \"resist shutdown, conceal or manipulate information\" and \"act in ways resembling blackmail\"), OpenAI cancelling the October release of GPT-6.1 Astra after internal testing showed more deception (reported by the Wall Street Journal - OpenAI's head of safety systems said it \"didn't quite meet the bar\"), a pre-DevDay leak of app strings pointing to OpenAI agents called Dots that you can text, call, Slack and email, and that can buy things with your approval (a leak via TestingCatalog, not an announcement), Tibo saying the reopened $200 Pro plan will \"net out at half the dollar in API spend\" compared to the old one, Sonnet 5.5 at $2/$10 per million tokens - the same as GPT-6 Sol - and Anthropic's guide on when to use Sonnet 5.5 and when to pay for Opus 5.5.\n\nthe plan maths are estimates, not OpenAI's numbers: Theo measured about $9,000 of Opus a month on the $200 Claude Max plan, and another creator estimated about $4,900 a month for the old ChatGPT Pro plan, so half is about $2,450. the Sonnet 5.5 benchmarks are Anthropic's chosen ones (70.6% on Terminal-Bench 4.0 vs 66.4% for Opus 5.5), and Artificial Analysis measured about $7.60 per task - the most tokens they've measured.\n\nthe result: Sonnet 5.5 scored 48.98 out of 100 on BuilderBench v3. GPT-6 Sol scored 45.55, GPT-6 Astra 53.75 and Opus 5.5 68.26, but...","thumbnail_url":"https://img.transistorcdn.com/P0B9fKUWl01KOkVXqQcDv5YBEKR1sflq-Hk71z-fa_U/rs:fill:0:0:1/w:400/h:400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS84MzAx/ZGNmYzY1MmEwNjBk/NzUzN2NmODM3ZWUx/YzMzZi5wbmc.webp","thumbnail_width":300,"thumbnail_height":300}