{"type":"rich","version":"1.0","provider_name":"Transistor","provider_url":"https://transistor.fm","author_name":"How I Tested That","title":"Teresa Torres | How I Tested an AI Coach","html":"<iframe width=\"100%\" height=\"180\" frameborder=\"no\" scrolling=\"no\" seamless src=\"https://share.transistor.fm/e/8354804d\"></iframe>","width":"100%","height":180,"duration":2936,"description":"Summary\n\nIn this episode, I’m joined by Teresa Torres. Teresa is an author, speaker and a product discovery coach who I’ve been a fan of for quite a long time. \nWe chat about why assumptions look different depending on the level you’re working at, why showing your work is so important for alignment, and how using AI inside your courses can give students more opportunities to practice and receive feedback.\nWe also dig into AI evals - how Teresa uses them to measure the quality of her AI coaching tools, identify failure modes, and systematically improve their performance instead of just trusting that an LLM will get it right on the first try.\nIf you want to learn more about testing assumptions, designing better learning loops, and using AI without giving up your critical thinking skills, this episode is for you.\n\nTakeaways\nAI can create a safer space for learning. Teresa found that people are often willing to ask an AI questions they might feel embarrassed asking another person.\nAn AI coach works best when it is grounded in a clear teaching model. Teresa’s interview coach was effective because it was built on years of refined curriculum, rubrics, and explicit feedback criteria.\nAI can dramatically increase opportunities for deliberate practice. Instead of waiting for an instructor, students can practice repeatedly and receive immediate, personalized feedback.\nBuilding AI tools can improve the underlying curriculum. When an agent struggles with ambiguous instructions, it exposes gaps in how the material itself is taught.\nEvals are simply a way to measure AI quality. Teresa uses evals to identify specific failure modes, track how often they occur, and test whether changes actually improve the agent.\nDomain expertise still matters. Recognizing that an AI coach has given subtly bad advice often requires deep knowledge of the subject, not just technical skill.\nProduct teams should define what “good” looks like for AI. Generic quality metrics are not enough; teams need...","thumbnail_url":"https://img.transistorcdn.com/hRAQ0Cvexq2Nhl7H1KPLfxWZ14skSKkH4xG8JMRnoOM/rs:fill:0:0:1/w:400/h:400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9zaG93/LzUwMDU0LzE3MDg3/MTI0NTQtYXJ0d29y/ay5qcGc.webp","thumbnail_width":300,"thumbnail_height":300}