{"type":"rich","version":"1.0","provider_name":"Transistor","provider_url":"https://transistor.fm","author_name":"TalkRL: The Reinforcement Learning Podcast","title":"Abhishek Naik on Continuing RL & Average Reward","html":"<iframe width=\"100%\" height=\"180\" frameborder=\"no\" scrolling=\"no\" seamless src=\"https://share.transistor.fm/e/778f5a6c\"></iframe>","width":"100%","height":180,"duration":4900,"description":"Abhishek Naik was a student at University of Alberta and Alberta Machine Intelligence Institute, and he just finished his PhD in reinforcement learning, working with Rich Sutton.  Now he is a postdoc fellow at the National Research Council of Canada, where he does AI research on Space applications. \nFeatured References \nReinforcement Learning for Continuing Problems Using Average Reward\nAbhishek Naik Ph.D. dissertation 2024 \nReward Centering\nAbhishek Naik, Yi Wan, Manan Tomar, Richard S. Sutton 2024   \nLearning and Planning in Average-Reward Markov Decision Processes\nYi Wan, Abhishek Naik, Richard S. Sutton 2020 \nDiscounted Reinforcement Learning Is Not an Optimization Problem \nAbhishek Naik, Roshan Shariff, Niko Yasui, Hengshuai Yao, Richard S. Sutton 2019  \n\nAdditional References \nExplaining dopamine through prediction errors and beyond, Gershman et al 2024 (proposes Differential-TD-like learning mechanism in the brain around Box 4)  ","thumbnail_url":"https://img.transistorcdn.com/jXB1-VPK-A9v1epzc4aG4pFxqlvo2vbQ_Ytyuar_gPI/rs:fill:0:0:1/w:400/h:400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9zaG93/LzIwNDcvMTcwNzk1/NDcxMS1hcnR3b3Jr/LmpwZw.webp","thumbnail_width":300,"thumbnail_height":300}