Find partners
TalkRL: The Reinforcement Learning Podcast

TalkRL: The Reinforcement Learning Podcast

Hosted by Robin Ranjit Singh Chauhan

TechnologyInterviews guests

Episodes

74

Latest episode

Aug 2026

Language

EN

About the show

TalkRL podcast is All Reinforcement Learning, All the Time. In-depth interviews with brilliant people at the forefront of RL research and practice. Guests from places like MILA, OpenAI, MIT, DeepMind, Berkeley, Amii, Oxford, Google Research, Brown, Waymo, Caltech, and Vector Institute. Hosted by Robin Ranjit Singh Chauhan.

Listen to episodes

60 recent
August 14, 2026Episode 741 hr 27 min

Thomas Frost on Clinical RL with Natural Timings

Dr Thomas Frost is an emergency physician based in London, UK. He is also in the final stages of completing a PhD at University College London, where he has been looking at offline reinforcement learning applied to healthcare settings. Featured References Robust Real-Time Mortality Prediction in the Intensive Care Unit using Temporal Difference Learning Thomas Frost, Kezhi Li, Steve Harris — ML4H Symposium, PMLR 259, 2025 Insulin4RL: Real-Time Insulin Infusions for Offline Reinforcement Learning Thomas Frost, Steve Harris — PhysioNet, 2026 (RRID:SCR_007345) The Hidden Risks of Temporal Resampling in Clinical Reinforcement Learning Thomas Frost, Hrisheekesh Vaidya, Steve Harris — arXiv preprint, 2026 Insulin4RL: Real-Time Insulin Management in the Intensive Care Unit for Offline Reinforcement Learning Thomas Frost, Steve Harris — arXiv preprint, 2026 Additional References The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care — Komorowski et al. 2018 Off by a beat: the effects of temporal misalignment in reinforcement learning for sepsis treatment — Tang et al. 2026 Identifying Decision Points for Safe and Interpretable Reinforcement Learning in Hypotension Treatment — Zhang et al. 2021 Where do doctors disagree? Characterizing Decision Points for Safe Reinforcement Learning in Choosing Vasopressor Treatment — Brown et al. 2025 Loss of plasticity in deep continual learning — Dohare et al. 2024

November 10, 2025Episode 731 hr 40 min

Danijar Hafner on Dreamer v4

Danijar Hafner was a Research Scientist at Google DeepMind until recently. Featured References Training Agents Inside of Scalable World Models [ blog ] Danijar Hafner, Wilson Yan, Timothy Lillicrap One Step Diffusion via Shortcut Models Kevin Frans, Danijar Hafner, Sergey Levine, Pieter Abbeel Action and Perception as Divergence Minimization [ blog ] Danijar Hafner, Pedro A. Ortega, Jimmy Ba, Thomas Parr, Karl Friston, Nicolas Heess Additional References Mastering Diverse Domains through World Models [ blog ] DreaverV3l Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, Timothy Lillicrap Mastering Atari with Discrete World Models [ blog ] DreaverV2 ; Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, Jimmy Ba Dream to Control: Learning Behaviors by Latent Imagination [ blog ] Dreamer ; Danijar Hafner, Timothy Lillicrap, Jimmy Ba, Mohammad Norouzi Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos [ Blog Post ], Baker et al

September 8, 2025Episode 7259 min

David Abel on the Science of Agency @ RLDM 2025

David Abel is a Senior Research Scientist at DeepMind on the Agency team, and an Honorary Fellow at the University of Edinburgh. His research blends computer science and philosophy, exploring foundational questions about reinforcement learning, definitions, and the nature of agency. Featured References Plasticity as the Mirror of Empowerment David Abel, Michael Bowling, André Barreto, Will Dabney, Shi Dong, Steven Hansen, Anna Harutyunyan, Khimya Khetarpal, Clare Lyle, Razvan Pascanu, Georgios Piliouras, Doina Precup, Jonathan Richens, Mark Rowland, Tom Schaul, Satinder Singh A Definition of Continual RL David Abel, André Barreto, Benjamin Van Roy, Doina Precup, Hado van Hasselt, Satinder Singh Agency is Frame-Dependent David Abel, André Barreto, Michael Bowling, Will Dabney, Shi Dong, Steven Hansen, Anna Harutyunyan, Khimya Khetarpal, Clare Lyle, Razvan Pascanu, Georgios Piliouras, Doina Precup, Jonathan Richens, Mark Rowland, Tom Schaul, Satinder Singh On the Expressivity of Markov Reward David Abel, Will Dabney, Anna Harutyunyan, Mark Ho, Michael Littman, Doina Precup, Satinder Singh — Outstanding Paper Award, NeurIPS 2021 Additional References Bidirectional Communication Theory — Marko 1973 Causality, Feedback and Directed Information — Massey 1990 The Big World Hypothesis — Javed et al. 2024 Loss of plasticity in deep continual learning — Dohare et al. 2024 Three Dogmas of Reinforcement Learning — Abel 2024 Explaining dopamine through prediction errors and beyond — Gershman et al. 2024 David Abel Google Scholar David Abel personal website

August 19, 2025Episode 7112 min

Jake Beck, Alex Goldie, & Cornelius Braun on Sutton's OaK, Metalearning, LLMs, Squirrels @ RLC 2025

Recorded at Reinforcement Learning Conference 2025 at University of Alberta, Edmonton Alberta Canada. Featured References Lecture on the Oak Architecture , Rich Sutton Alberta Plan , Rich Sutton with Mike Bowling and Patrick Pilarski Additional References Jacob Beck on Google Scholar Alex Goldie on Google Scholar Cornelius Braun on Google Scholar Reinforcement Learning Conference

August 18, 2025Episode 7014 min

Outstanding Paper Award Winners - 2/2 @ RLC 2025

We caught up with the RLC Outstanding Paper award winners for your listening pleasure. Recorded on location at Reinforcement Learning Conference 2025 , at University of Alberta, in Edmonton Alberta Canada in August 2025. Featured References Empirical Reinforcement Learning Research Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions Ayush Jain, Norio Kosaka, Xinhu Li, Kyung-Min Kim, Erdem Biyik, Joseph J Lim Applications of Reinforcement Learning WOFOSTGym: A Crop Simulator for Learning Annual and Perennial Crop Management Strategies William Solow, Sandhya Saisubramanian, Alan Fern Emerging Topics in Reinforcement Learning Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners Calarina Muslimani, Kerrick Johnstonbaugh, Suyog Chandramouli, Serena Booth, W. Bradley Knox, Matthew E. Taylor Scientific Understanding in Reinforcement Learning Multi-Task Reinforcement Learning Enables Parameter Scaling Reginald McLean, Evangelos Chatzaroulas, J K Terry, Isaac Woungang, Nariman Farsad, Pablo Samuel Castro

August 15, 2025Episode 696 min

Outstanding Paper Award Winners - 1/2 @ RLC 2025

We caught up with the RLC Outstanding Paper award winners for your listening pleasure. Recorded on location at Reinforcement Learning Conference 2025 , at University of Alberta, in Edmonton Alberta Canada in August 2025. Featured References Scientific Understanding in Reinforcement Learning How Should We Meta-Learn Reinforcement Learning Algorithms? Alexander David Goldie, Zilin Wang, Jakob Nicolaus Foerster, Shimon Whiteson Tooling, Environments, and Evaluation for Reinforcement Learning Syllabus: Portable Curricula for Reinforcement Learning Agents Ryan Sullivan, Ryan Pégoud, Ameen Ur Rehman, Xinchen Yang, Junyun Huang, Aayush Verma, Nistha Mitra, John P Dickerson Resourcefulness in Reinforcement Learning PufferLib 2.0: Reinforcement Learning at 1M steps/s Joseph Suarez Theory of Reinforcement Learning Deep Reinforcement Learning with Gradient Eligibility Traces Esraa Elelimy, Brett Daley, Andrew Patterson, Marlos C. Machado, Adam White, Martha White

August 4, 2025Episode 6852 min

Thomas Akam on Model-based RL in the Brain

Prof Thomas Akam is a Neuroscientist at the Oxford University Department of Experimental Psychology. He is a Wellcome Career Development Fellow and Associate Professor at the University of Oxford, and leads the Cognitive Circuits research group . Featured References Brain Architecture for Adaptive Behaviour Thomas Akam, RLDM 2025 Tutorial Additional References Thomas Akam on Google Scholar pyPhotometry : Open source, Python based, fiber photometry data acquisition pyControl : Open source, Python based, behavioural experiment control. Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control , Nathaniel D Daw, Yael Niv, Peter Dayan, 2005 Further analysis of the hippocampal amnesic syndrome: 14-year follow-up study of H. M. , Milner, B., Corkin, S., & Teuber, H. L., 1968 Internally generated cell assembly sequences in the rat hippocampus , Pastalkova E, Itskov V, Amarasingham A, Buzsáki G. Science. 2008 Multi-disciplinary Conference on Reinforcement Learning and Decision 2025

July 22, 2025Episode 6731 min

Stefano Albrecht on Multi-Agent RL @ RLDM 2025

Stefano V. Albrecht was previously Associate Professor at the University of Edinburgh, and is currently serving as Director of AI at startup Deepflow . He is a Program Chair of RLDM 2025 and is co-author of the MIT Press textbook " Multi-Agent Reinforcement Learning: Foundations and Modern Approaches ". Featured References Multi-Agent Reinforcement Learning: Foundations and Modern Approaches Stefano V. Albrecht, Filippos Christianos, Lukas Schäfer MIT Press, 2024 RLDM 2025: Reinforcement Learning and Decision Making Conference Dublin, Ireland EPyMARL: Extended Python MARL framework https://github.com/uoe-agents/epymarl Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms in Cooperative Tasks Georgios Papoudakis and Filippos Christianos and Lukas Schäfer and Stefano V. Albrecht

June 25, 2025Episode 665 min

Satinder Singh: The Origin Story of RLDM @ RLDM 2025

Professor Satinder Singh of Google DeepMind and U of Michigan is co-founder of RLDM. Here he narrates the origin story of the Reinforcement Learning and Decision Making meeting (not conference). Recorded on location at Trinity College Dublin, Ireland during RLDM 2025. Featured References RLDM 2025: Multi-disciplinary Conference on Reinforcement Learning and Decision Making (RLDM) June 11-14, 2025 at Trinity College Dublin, Ireland Satinder Singh on Google Scholar

March 9, 2025Episode 6510 min

NeurIPS 2024 - Posters and Hallways 3

Posters and Hallway episodes are short interviews and poster summaries. Recorded at NeurIPS 2024 in Vancouver BC Canada. Featuring Claire Bizon Monroc from Inria: WFCRL: A Multi-Agent Reinforcement Learning Benchmark for Wind Farm Control Andrew Wagenmaker from UC Berkeley: Overcoming the Sim-to-Real Gap: Leveraging Simulation to Learn to Explore for Real-World RL Harley Wiltzer from MILA: Foundations of Multivariate Distributional Reinforcement Learning Vinzenz Thoma from ETH AI Center: Contextual Bilevel Reinforcement Learning for Incentive Alignment Haozhe (Tony) Chen & Ang (Leon) Li from Columbia: QGym: Scalable Simulation and Benchmarking of Queuing Network Controllers

Is this your show?

Claim this listing to keep it up to date, reach guests who want to pitch you, and manage bookings with Guestify.

Claim this listing

More Technology podcasts