{"type":"rich","version":"1.0","provider_name":"Transistor","provider_url":"https://transistor.fm","author_name":"The Experimentation Edge","title":"How Fin A/B Tests Millions of Samples in Days","html":"<iframe width=\"100%\" height=\"180\" frameborder=\"no\" scrolling=\"no\" seamless src=\"https://share.transistor.fm/e/0dd7272a\"></iframe>","width":"100%","height":180,"duration":2734,"description":"SummaryOn this episode of The Experimentation Edge, host Ashley Stirrup talks with Pedro Tabacof, Principal Machine Learning Scientist at Fin (formerly Intercom), about how one of the most advanced AI customer support agents in the world is built on relentless experimentation. Pedro explains why unit tests don't work on non-deterministic AI, how Fin runs up to two dozen concurrent A/B tests pulling millions of samples in days, and shares two counterintuitive experiments: one where slowing the agent down improved every metric, and one where adding more context made Fin more helpful and more prone to fake promises until a targeted prompt fix kept the upside without the hallucinations. It's a candid look for product managers, engineers, and data scientists at how a $100M ARR AI product actually ships improvements.\nChapters00:00 Welcome and what Fin actually does02:00 How Fin became Anthropic's first line of support02:30 Why Fin sells resolutions not deflections06:00 Owning the stack with custom models10:40 Pedro's path from fuzzy logic to AI12:55 Why A/B testing is the only gold standard for AI15:50 Do no harm testing on every change18:00 The latency experiment that shocked the team27:30 When more context made Fin hallucinate30:15 Win rates and the future of AI driven experimentation\nTakeaways-Faster is not always better. Fin increased latency artificially and positive feedback went up, likely because a small delay makes an AI feel like it is doing real work.-You cannot unit test a non-deterministic AI. A/B testing at scale, millions of samples in days, is the only reliable way to know a change actually helped.-Adding more conversation history made Fin more helpful and more prone to fake promises, until a targeted prompt fix removed the hallucinations and kept most of the gain.-A losing experiment is often a winner with one broken part. Diagnose which element hurts the experience, fix only that, and rerun.-Fin A/B tests everything, even one-character prompt changes...","thumbnail_url":"https://img.transistorcdn.com/D9kLs0HSsqR4ttk_5ESEdC1jX-wmD76GK-OHmb3a9B8/rs:fill:0:0:1/w:400/h:400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS80YTFk/MGU1MjJlODhlNjJh/MTdlZTZkN2Q1ODY5/OTdjYy5wbmc.webp","thumbnail_width":300,"thumbnail_height":300}