{"type":"rich","version":"1.0","provider_name":"Transistor","provider_url":"https://transistor.fm","author_name":"TalkRL: The Reinforcement Learning Podcast","title":"Ian Osband","html":"<iframe width=\"100%\" height=\"180\" frameborder=\"no\" scrolling=\"no\" seamless src=\"https://share.transistor.fm/e/f818c3cb\"></iframe>","width":"100%","height":180,"duration":4106,"description":"Ian Osband is a Research scientist at OpenAI (ex DeepMind, Stanford) working on decision making under uncertainty.  \nWe spoke about: \n- Information theory and RL \n- Exploration, epistemic uncertainty and joint predictions \n- Epistemic Neural Networks and scaling to LLMs \n\nFeatured References \nReinforcement Learning, Bit by Bit \nXiuyuan Lu, Benjamin Van Roy, Vikranth Dwaracherla, Morteza Ibrahimi, Ian Osband, Zheng Wen \nFrom Predictions to Decisions: The Importance of Joint Predictive Distributions \nZheng Wen, Ian Osband, Chao Qin, Xiuyuan Lu, Morteza Ibrahimi, Vikranth Dwaracherla, Mohammad Asghari, Benjamin Van Roy  \n \nEpistemic Neural Networks \nIan Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, Benjamin Van Roy  \n\nApproximate Thompson Sampling via Epistemic Neural Networks \nIan Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, Benjamin Van Roy \n  \n\nAdditional References  \nThesis defence, Ian Osband \nHomepage, Ian Osband \nEpistemic Neural Networks at Stanford RL Forum \nBehaviour Suite for Reinforcement Learning, Osband et al 2019 \nEfficient Exploration for LLMs, Dwaracherla et al 2024 ","thumbnail_url":"https://img.transistorcdn.com/jXB1-VPK-A9v1epzc4aG4pFxqlvo2vbQ_Ytyuar_gPI/rs:fill:0:0:1/w:400/h:400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9zaG93/LzIwNDcvMTcwNzk1/NDcxMS1hcnR3b3Jr/LmpwZw.webp","thumbnail_width":300,"thumbnail_height":300}