Future of Life Institute Podcast

David Manheim is head of methodology at AI Evaluation Consensus. He joins the podcast to discuss how AI evaluations can become more reliable, transparent, and useful for decisions. We cover common failures such as unclear reporting, training to the test, benchmark saturation, and models changing behavior when they know they are being tested. The conversation also examines real-world tests, biosecurity, persuasion, forecasting, human oversight, and why even “normal” AI progress could be disruptive.


LINKS:

CHAPTERS:
(00:00) Episode Preview
(01:04) Evaluation consensus project
(07:01) Evaluation awareness challenges
(12:28) Reporting capabilities clearly
(19:38) Benchmarks beyond humans
(29:52) Proxies and biosecurity
(42:01) Persuasion and democracy
(53:59) Forecasting with AI
(01:08:44) Oversight and disruption
(01:16:42) Supporting better evals


PRODUCED BY:
https://aipodcast.ing


SOCIAL LINKS:
Website: https://podcast.futureoflife.org
Twitter (FLI): https://x.com/FLI_org
Twitter (Gus): https://x.com/gusdocker
LinkedIn: https://www.linkedin.com/company/future-of-life-institute/
YouTube: https://www.youtube.com/channel/UC-rCCy3FQ-GItDimSR9lhzw/
Apple: https://geo.itunes.apple.com/us/podcast/id1170991978
Spotify: https://open.spotify.com/show/2Op1WO3gwVwCrYHg4eoGyP


What is Future of Life Institute Podcast?

The Future of Life Institute (FLI) is a nonprofit working to reduce global catastrophic and existential risk from powerful technologies. In particular, FLI focuses on risks from artificial intelligence (AI), biotechnology, nuclear weapons and climate change. The Institute's work is made up of three main strands: grantmaking for risk reduction, educational outreach, and advocacy within the United Nations, US government and European Union institutions. FLI has become one of the world's leading voices on the governance of AI having created one of the earliest and most influential sets of governance principles: the Asilomar AI Principles.