AI News Today | Julian Goldie Podcast

Claude Sonnet 5 Review: More Expensive, Worse Than Opus 4.8? (Benchmarks & Agent Tests)

The video reviews Anthropic’s newly released Claude Sonnet 5, described as more agentic and capable of planning and tool use, but argues it underperforms Opus 4.8 on benchmarks (including agentic coding) while costing more. The creator shares Goldy Bench examples Sonnet 5 generated (a ray caster maze, a broken galaxy orbit test, a synthwave background, and a crypt game), noting some outputs look good but others fail. Side-by-side comparisons show mixed results versus GLM 5.2, with GLM succeeding on tasks Sonnet 5 fails, and tweets highlight negative reception focused on poor token efficiency and pricing. The recommendation is to keep using Opus 4.8, expect Fable 5 soon, and focus on building flexible agent systems that can swap models in and out.

00:00 Sonnet 5 Launch
00:30 Benchmarks vs Opus
01:39 Goldy Bench Demos
02:53 GLM 5.2 Comparisons
04:00 Backlash and Pricing
05:57 Fugu Ultra Showdown
07:20 Why Release This
08:00 Focus on Systems
09:11 Agent OS Pitch
09:48 Final Verdict

Creators and Guests

Host
Julian Goldie
Founder of AI Profit Boardroom and your daily guide to the AI revolution. I break down the biggest AI news, agent updates, and breakthroughs — fast, clear, and no hype.

What is AI News Today | Julian Goldie Podcast?

Latest Podcast