{"type":"rich","version":"1.0","provider_name":"Transistor","provider_url":"https://transistor.fm","author_name":"TalkRL: The Reinforcement Learning Podcast","title":"Natasha Jaques 2","html":"<iframe width=\"100%\" height=\"180\" frameborder=\"no\" scrolling=\"no\" seamless src=\"https://share.transistor.fm/e/a18817da\"></iframe>","width":"100%","height":180,"duration":2762,"description":"Hear about why OpenAI cites her work in RLHF and dialog models, approaches to rewards in RLHF, ChatGPT, Industry vs Academia, PsiPhi-Learning, AGI and more! \nDr Natasha Jaques is a Senior Research Scientist at Google Brain.\nFeatured References\nWay Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog\nNatasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, Rosalind Picard \n\nSequence Tutor: Conservative Fine-Tuning of Sequence Generation Models with KL-control\nNatasha Jaques, Shixiang Gu, Dzmitry Bahdanau, José Miguel Hernández-Lobato, Richard E. Turner, Douglas Eck \n\nPsiPhi-Learning: Reinforcement Learning with Demonstrations using Successor Features and Inverse Temporal Difference Learning\nAngelos Filos, Clare Lyle, Yarin Gal, Sergey Levine, Natasha Jaques, Gregory Farquhar  \nBasis for Intentions: Efficient Inverse Reinforcement Learning using Past Experience\nMarwa Abdulhai, Natasha Jaques, Sergey Levine \nAdditional References  \nFine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al 2019  \nLearning to summarize from human feedback, Nisan Stiennon et al 2020  \nTraining language models to follow instructions with human feedback, Long Ouyang et al 2022  ","thumbnail_url":"https://img.transistorcdn.com/jXB1-VPK-A9v1epzc4aG4pFxqlvo2vbQ_Ytyuar_gPI/rs:fill:0:0:1/w:400/h:400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9zaG93/LzIwNDcvMTcwNzk1/NDcxMS1hcnR3b3Jr/LmpwZw.webp","thumbnail_width":300,"thumbnail_height":300}