{"type":"rich","version":"1.0","provider_name":"Transistor","provider_url":"https://transistor.fm","author_name":"Thinking Machines: AI & Philosophy","title":"LLM Inference Speed (Tech Deep Dive)","html":"<iframe width=\"100%\" height=\"180\" frameborder=\"no\" scrolling=\"no\" seamless src=\"https://share.transistor.fm/e/e71cd65d\"></iframe>","width":"100%","height":180,"duration":2376,"description":"In this tech talk, we dive deep into the technical specifics around LLM inference.\nThe big question is: Why are LLMs slow? How can they be faster? And might slow inference affect UX in the next generation of AI-powered software?\n\nWe jump into:\nIs fast model inference the real moat for LLM companies?\nWhat are the implications of slow model inference on the future of decentralized and edge model inference?\nAs demand rises, what will the latency/throughput tradeoff look like?\nWhat innovations on the horizon might massively speed up model inference?","thumbnail_url":"https://img.transistorcdn.com/S6OjXZjcpOAZ6jDX4fc4XvtZpoBLPUwb1-xRPS1F5K0/rs:fill:0:0:1/w:400/h:400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9zaG93/LzQ1MTk5LzE3MDg3/MDIxODItYXJ0d29y/ay5qcGc.webp","thumbnail_width":300,"thumbnail_height":300}