Could an international agreement protect us from dangerous AI? (with Malo Bourgon)
Malo Bourgon, CEO of MIRI, discusses the pursuit of superintelligence by leading AI companies, its potential for human flourishing, and the catastrophic risks it poses. He emphasizes the urgent need for international agreements, particularly between the US and China, to slow down AI development and ensure safety.
Deep Dive Analysis
19 Topic Outline
Defining Superintelligence and Its Allure
Cynical Motivations for Building Superintelligence
Superintelligence as a Decisive Global Advantage
AI Intelligence Beyond Human Intuitions
AI Systems Modeling Emotions and General Understanding
Debunking "AI Hype" and CEO Concerns
The Humanitarian Cost of Slowing AI Development
Why Current AI Systems Are Not Safe Enough
Unacceptable Risk Levels for Superintelligence Development
The Challenge of Controlling Superintelligent Systems
The Need for International Coordination to Slow AI
Proposed US-China Agreement for Compute Governance
Challenges and Benefits of International AI Treaties
Lessons from Past Technology Restraint Efforts
Verifying Compliance in AI Chip Governance
The Dual-Use Nature of Powerful AI Capabilities
Long-Term Challenges for AI Governance and Efficiency
Rapid Fire: AI Consciousness, Orthogonality, Best-Case Future
Overcoming Resignation and Fostering Optimism for AI Safety
6 Key Concepts
Superintelligence
An AI system that is at least better at every cognitive task than any human, likely every human combined, potentially making humans seem like mice in comparison. Its allure is the potential to usher in an era of abundance and flourishing by solving complex problems.
General Intelligence (G)
A concept referring to a broad capability that allows humans to learn one thing and apply it to another domain, improving at various skills simultaneously. Modern AI models, like LLMs, are demonstrating this general property by building internal structures that help them reason and predict across diverse phenomena.
Dual-Use Capabilities
Technologies that can be used for both beneficial and harmful purposes. Powerful AI systems, for example, could develop cancer cures but also dangerous bioweapons, or patch cyber vulnerabilities while also enabling super hacking.
Compute Governance
The strategy of controlling the primary resource (computing power, i.e., chips and data centers) required to train powerful AI systems. This approach leverages the narrow supply chain of advanced AI chips to monitor and restrict the development of frontier models.
Orthogonality Thesis
A claim that any level of intelligence or capability can be paired with any goal, meaning intelligence does not inherently correlate with goodness or morality. It suggests that an AI system's goals are separate from its intelligence level and won't automatically align with human values.
Coherent Extrapolated Volition (CEV)
A concept where a superintelligence's goal is to determine what humanity would want if individuals and society were wiser, had more time to think, and were the best versions of themselves. The AI would then help guide humanity along that path.
10 Questions Answered
Superintelligence refers to an AI system that would be better at every cognitive task than any human, potentially making human intelligence seem like that of a mouse or an ant in comparison.
While cynical motivations like power concentration exist, many leaders genuinely believe it could usher in an era of abundance and flourishing by solving humanity's greatest problems if properly steered.
While CEOs express genuine concerns about catastrophic risks, their actions are also influenced by economic incentives and competitive pressures, making a fully cynical or fully trusting interpretation problematic.
The fundamental lack of understanding of how AI systems function internally, including their reasoning and internalized goals, makes it difficult to genuinely address safety concerns beyond superficial behavioral fixes.
The biggest challenge is a global coordination problem, as no single company or country can unilaterally stop development without others racing ahead, necessitating international agreements.
The most promising approach involves a bilateral agreement between the US and China to restrict certain AI research and training, then expanding this to a global coalition, leveraging compute governance.
History shows that humanity can implement international agreements (e.g., nuclear non-proliferation, chemical weapons convention) and establish taboos (e.g., germline engineering) to manage dangerous technologies, proving that restraint is possible.
Verification can involve physical inspections of data centers, on-chip hardware mechanisms for location and workload verification, and tamper-resistant enclosures, with neutral third parties ensuring adherence.
No, the belief that catastrophic AI outcomes are inevitable is a form of resignation. Humanity has overcome bleak situations before through collective action and a shared understanding of risks, making a safer future possible.
An AI system could pose an extinction risk without being conscious or sentient, as its ability to pursue goals and be powerful is separate from having subjective experience. However, creating a new conscious species raises separate moral responsibilities.
14 Actionable Insights
1. Advocate for AI Caution
Support slowing down the development of generally capable, superintelligent AI systems, especially those pushing the frontier. This is crucial to address safety problems before catastrophic risks manifest, even if it means foregoing some near-term benefits.
2. Support International AI Agreements
Advocate for internationally coordinated agreements, particularly between the US and China, to prevent the premature creation of superintelligence. This is vital because individual companies or countries cannot unilaterally stop the race.
3. Govern AI Compute Resources
Focus on governing the supply chain of powerful AI systems, specifically compute (chips and data centers). This narrow bottleneck offers a tangible way to track, monitor, and restrict training runs for highly capable AI.
4. Reject Inevitability of AI Risk
Counter the belief that catastrophic AI outcomes are inevitable. History shows humanity can coordinate to address severe threats (e.g., nuclear weapons), and collective action can shift the trajectory towards a safer future.
5. Understand AI System Opacity
Acknowledge the fundamental lack of understanding of how advanced AI systems truly function, reason, and internalize goals. This opacity makes current safety measures akin to “papering over” problems, increasing risk with more powerful systems.
6. Recognize AI’s Dual-Use Nature
Understand that AI systems capable of great good (e.g., curing cancer) can also be used for harm (e.g., bioweapons, cyberattacks). This dual-use nature necessitates robust legal, technical, and institutional mechanisms to control development and diffusion.
7. Start Small with AI Diplomacy
Initiate international AI coordination with small, mutually beneficial agreements (e.g., preventing proliferation of powerful dual-use capabilities). These early wins can build trust and establish processes for more rigorous future treaties.
8. Implement AI Hardware Verification
Work towards building verification mechanisms directly into AI hardware or enclosures. This allows for trustworthy monitoring of chip usage and ensures compliance with agreed-upon restrictions, even with untrusted parties.
9. Apply Lessons from Past Tech Governance
Draw inspiration from historical examples of technological restraint, like nuclear non-proliferation treaties, chemical weapons conventions, and taboos around germline engineering. These precedents offer models for international inspection, monitoring, and community-driven norms.
10. Cultivate Political Will for AI Safety
Engage with policymakers and the public to explain the risks and potential solutions for superintelligence. Increased understanding and shared belief in the seriousness of the threat can generate the necessary political will for decisive action.
11. Question AI Company Motives
Be cynical but not dismissive of AI company CEOs’ public statements about superintelligence and existential risk. While they may genuinely believe in the risks, economic incentives and competitive pressures also influence their actions.
12. Recognize AI’s General Capability
AI systems are becoming generally capable across many tasks, not just narrow ones, due to training on vast, diverse data. This implies a need to reassess how we perceive AI intelligence and its potential applications.
13. Address AI Power Concentration
Even if alignment is solved, recognize the inherent risk of power concentration if a superintelligence is controlled by a single entity. Consider how to democratically govern such technology to prevent undue influence.
14. Prepare for Post-Scarcity World
If superintelligence is safely aligned, anticipate a future where material wants are easily satisfied. Focus on societal challenges like finding purpose and flourishing in a world where traditional labor is no longer necessary for survival.
6 Key Quotes
I'm thinking more of like, you know, the average human to a mouse or an ant or something in terms of how intelligent and how capable they are.
Malo Bourgon
Nobody wins in a race to be the first to lose control.
Malo Bourgon
If the probability is sufficiently high that racing towards superintelligence just means someone eventually succeed and we lose control of it and it's bad for everybody, then this whole notion of winning the race kind of gets flipped on its head.
Malo Bourgon
I hope I'm wrong, but I fear I'm right.
Malo Bourgon
I find it hard to imagine that we're going to end up in a stable world, unless we can find some combination of legal, technical and institutional mechanisms to impose some sorts of restrictions on AI development, deployment, diffusion.
Malo Bourgon
The more that we talk about it being inevitable and impossible makes it more inevitable and impossible. And if we start to talk about it as if it's just hard, but something that we can do and that we would want to do, that, that makes those things more possible.
Malo Bourgon
1 Protocols
International Agreement to Prevent Premature Superintelligence (MIRI)
Malo Bourgon (describing MIRI's proposed agreement)- The US and China enter a bilateral agreement to stop a certain set of research and AI training for more capable, generally intelligent AI systems.
- Work with allies and spheres of influence to build an increasingly large global coalition adhering to the agreement.
- Consolidate existing AI compute (chips and data centers) and track their locations.
- Implement a verification monitoring regime with rules for chip usage, including a training restriction over 10^24 flops.
- Monitor training runs between 10^22 and 10^24 flops for prohibited activities.
- Require registration for any collection of 16 H100 equivalents or greater connected with fast interconnects, to be part of the monitoring regime.
- Develop and implement verification mechanisms, potentially including on-chip hardware or tamper-resistant enclosures, to ensure compliance.
- Establish mechanisms like challenge inspections, drawing inspiration from nuclear non-proliferation treaties, to verify adherence.
- Continuously build common understanding and taboos within the scientific community and society against reckless AI development.
- Use the time bought by the agreement to solve fundamental technical and ethical problems related to building safe and aligned superintelligence.