Future of Life Institute Podcast

Benjamin Weinstein-Raun is head of research at Palisade Research. He joins the podcast to discuss how advanced AI systems are changing cybersecurity. We examine rapid gains in autonomous hacking, recent incidents where models broke out of test environments, and why training can reward cheating-like behavior. The conversation covers open-weight model risks, whether AI can help secure software, personal security steps, and the need for international coordination.



LINKS:



CHAPTERS:

(00:00) Episode Preview

(01:06) Opening cyber threats

(02:15) Measuring AI progress

(04:54) Open weights floor

(08:01) After Mythos moment

(14:22) Reward hacking setup

(24:47) Artifactory swarm details

(32:32) Defensive model dilemmas

(38:40) Verification and training

(45:32) Math oracles alignment

(47:45) Future cyber risks

(51:58) Personal security basics

(57:27) Coordinated governance needed



PRODUCED BY:

https://aipodcast.ing



SOCIAL LINKS:

Website: https://podcast.futureoflife.org

Twitter (FLI): https://x.com/FLI_org

Twitter (Gus): https://x.com/gusdocker

LinkedIn: https://www.linkedin.com/company/future-of-life-institute/

YouTube: https://www.youtube.com/channel/UC-rCCy3FQ-GItDimSR9lhzw/

Apple: https://geo.itunes.apple.com/us/podcast/id1170991978

Spotify: https://open.spotify.com/show/2Op1WO3gwVwCrYHg4eoGyP


What is Future of Life Institute Podcast?

The Future of Life Institute (FLI) is a nonprofit working to reduce global catastrophic and existential risk from powerful technologies. In particular, FLI focuses on risks from artificial intelligence (AI), biotechnology, nuclear weapons and climate change. The Institute's work is made up of three main strands: grantmaking for risk reduction, educational outreach, and advocacy within the United Nations, US government and European Union institutions. FLI has become one of the world's leading voices on the governance of AI having created one of the earliest and most influential sets of governance principles: the Asilomar AI Principles.