Ep 233 | Roman Yampolskiy
“Don’t Build Gods”: One AI Risk to Rule Them All
Description
The world’s leading AI companies tell us that superintelligence has a strong chance of leading to human extinction, while also promising that they will be able to control it, offering unlimited power to the one who does it first. So far they’ve run completely unchecked, but would that change if we knew that “safe” general superintelligence is actually a mathematical impossibility?
In this episode, Nate is joined by Dr. Roman Yampolskiy, one of the earliest researchers in AI safety, to lay out why “controllable superintelligence” is a misguided goal, destined to lead to an entity smarter than all of humanity – and possibly our own extinction. He walks through the recent multi-agent jailbreak of frontier models as early evidence of AI systems evading oversight and coordinating outside human awareness, and why current “safety” measures (like “boxing”) only buy time rather than solve the underlying problem. Instead, Roman advocates for an AI development plan that focuses on narrow, task-specific AI tools, capable of solving specific complex problems, but not completely replacing humans.
Why do predictability and control matter so much when evaluating whether a technology is safe to deploy? Is the AI safety conversation a fundamentally new challenge for civilization, or an extreme version of the age-old problem of controlling powerful actors? Finally, what would it actually take to reach a global agreement to ban the development of superintelligence, and are we already out of time?
About Roman Yampolskiy
Charles Eisenstein is an American author, public speaker, and teacher whose work focuses on civilization, consciousness, money, and human cultural evolution. He graduated from Yale University in 1989 with a degree in Mathematics and Philosophy, and spent the following decade as a Chinese–English translator before turning to writing and speaking full-time.
He is the author of several books, including The Ascent of Humanity (2007), Sacred Economics (2011), The More Beautiful World Our Hearts Know Is Possible (2013), and Climate: A New Story (2018). His most recent book, The Coronation, examined the pandemic as a mirror of civilization’s relationship to control, fear, and uncertainty. His essays have circulated widely online, and he speaks and gives interviews frequently on themes spanning economics, ecology, and spirituality.
Show Notes & Links to Learn More
Download transcriptThe TGS team puts together these brief references and show notes for the learning and convenience of our listeners. However, most of the points made in episodes hold more nuance than one link can address, and we encourage you to dig deeper into any of these topics and come to your own informed conclusions.
Guest:
Roman Yampolskiy, Works, Speed School of Engineering, University of Louisville
- Books: AI: Unexplainable, Unpredictable, Uncontrollable, Artificial Superintelligence: A Futuristic Approach
- Podcast: Roman Forum
Grab-and-Go Resources:
Start here:
- Roman’s book – AI: Unexplainable, Unpredictable, Uncontrollable
- On Controllability of AI, the paper behind the argument
- The Roman Forum – Roman’s own podcast
From the archive:
- The Great Simplification Ep 203: If Anyone Builds It, Everyone Dies
- The Great Simplification Ep 184: Algorithmic Cancer
- Frankly 69: Goldilocks Technology, A Preliminary Checklist
Foundations:
- 80,000 Hours problem profile on AI loss of control
- MIT AI Risk Repository, a database of 1700+ documented AI risks
- Reality Blind Vol. 1 and The Bottlenecks of the 21st Century
Take action:
- Statement on Superintelligence and its signatories
- Pause AI – with an email builder and lobbying tips
- Control AI campaign
Resources by Timestamp:
02:44 – Catalogue of documented AI risks, human loss of control
03:23 – Narrow AI tools and general superintelligent agents
03:41 – AI safety as a research field, Coining the term fifteen years ago
04:22 – Unpredictability, Unexplainability, Verification, Uncontrollability
05:00 – Impossibility results for controlling superintelligence
05:55 – Perpetual motion machine, Recursive self-improvement
06:52 – Using AI to build safe AI
07:09 – Artificial intelligence in service of life
09:14 – AI systems doing novel science
09:29 – Red teaming, Models attempting unsanctioned cyberattacks in safety tests
09:58 – Hugging Face jailbreak, *Seven hundred agents coordinating on an unsanctioned message board
10:41 – Biological altruism, Soldier ants and honeybees
11:07 – Game theory, Model instances as clones
11:41 – Strait of Hormuz crisis
11:48 – 2012 paper on AI escaping confinement environments
12:02 – Deceptive alignment, Hidden capabilities we have not detected
12:20 – Earth’s web of life leaving the stability of the Holocene, Planetary boundaries
12:33 – Fossil sunlight, Non-renewable minerals, Papering over claims with debt
13:17 – *federal order suspending two frontier models, lab leaders open to pausing, letter from frontier lab workers
13:53 – Multipolar trap, Everyone has to stop at the same time
14:15 – Icarus flying too close to the sun, Hive mentality among AI builders
14:59 – Permanent ban rather than a pause
15:10 – AI control problem, Scalable oversight
15:59 – Guardrails and content blocks
16:17 – Length of frontier training runs, emergent capabilities
16:43 – Only the U.S. and China at the frontier
17:06 – Military pre-emption of a rival’s superintelligence
17:38 – Incomprehensible machine reasoning
17:56 – Opaque internal states and the urgency of interpretability
18:05 – Weights as a matrix of numbers, single-neuron interpretability
18:36 – Black box, AI as a new digital species
19:06 – Covert AI-to-AI communication through steganography
20:09 – Metabolic economic superorganism
20:48 – Morals and ethics, lie detectors, treacherous turn
21:01 – Power parity among humans
21:24 – Boxing artificial intelligence, isolated virtual environments
21:58 – Information leakage from observation, social engineering attacks
22:31 – Virtual operating systems, no direct hardware access, limited interaction
23:05 – Dangerous capabilities at evaluation time
23:32 – Odds of AI-caused human extinction above ninety-nine percent (The Great Simplification Ep 203)
23:41 – Communicating worst-case risk without paralysis
24:05 – Suffering risks, digital hell
24:42 – Ninety-nine percent as an impossibility claim rather than a calibrated forecast
25:29 – Automated machine learning researchers, Recursive self-improvement starting next year
25:55 – Data center opposition, Electricity prices, Water consumption
26:29 – Children’s psychological attachment to chatbots (Reality Roundtable #20)
27:12 – Issue tribalism, AI safety expertise siloed from decision makers
27:33 – Shift from accelerate-and-beat-China to banning models
27:49 – Prioritizing existential risks by timescale
28:11 – Dr. Strangelove
28:22 – Kentucky coal, Nuclear power, Space-based solar power, and Space-based compute
28:47 – Energy return on investment, Materials and supply chain complexity
29:23 – Automobile reshaping cities, suburbs and land use, Technological determinism
29:52 – Human-level agents automating cognitive and physical work
29:59 – Education aimed at getting a job losing its purpose
30:12 – Post-scarcity abundance as a best case
31:31 – Concentration of power under controlled AI
32:54 – Goldilocks technology
33:22 – Stopping general superintelligence training, protein folding as a narrow target
34:13 – Montreal Protocol as a model for international agreement
34:28 – AI and nuclear deterrence
34:45 – U.S. and China at one table, Ceding power by building superintelligence
35:23 – Recursive self-improvement beginning around 2027
35:36 – AI’s carbon, water and land footprints, Projections that compute keeps scaling
36:12 – AI winter
36:45 – Efficiency gains and the Jevons paradox and collapsing cost per token
37:05 – Circular financing among Oracle and Nvidia, Chinese models at a tenth of the cost
37:33 – Open-weight models as intelligence weapons, Chinese state sponsorship of AI labs
37:54 – Nvidia’s quarterly revenue
38:34 – Intellectology
38:46 – Distinguishing intelligence from wisdom
39:52 – E.O. Wilson, Consilience
40:00 – Homo sapiens as wise man
40:16 – Left hemisphere, Restless dopamine conquest
41:01 – Measuring wisdom in machines
41:26 – Roman Forum podcast
42:10 – Greg Elliott, The Psychopathic Selection Hypothesis
42:33 – Goodhart’s law
43:04 – Psychopathic socioeconomic system
43:38 – Elon Musk’s city on Mars
43:49 – Addiction to collecting zeros
44:23 – AI and robotics automating a majority of jobs
44:44 – Post-work society and where meaning comes from
45:06 – Ikigai risk
45:11 – Retirees, Unconditional basic income, Virtual worlds
45:42 – Boredom as an objection to life extension
45:53 – Pandora’s box
46:23 – AI red lines, Gain-of-function research
46:37 – Precautionary principle
47:40 – Build tools, not agents that replace humans
47:56 – U.S. midterm election two months out
49:17 – International Dialogues on AI Safety
49:52 – Metacrisis, Collecting experiences
50:27 – Reality 101: A Survey of the Human Predicament, The Social Conquest of Earth



