What Dario Amodei thinks about AI alignment
Co-founder and CEO of Anthropic; previously VP of research at OpenAI.
Everything they publish, on ppll ↗
Dario Amodei did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
2 dated positions, 2025 to 2026, in their own words. Our reading of what Dario Amodei has said — not written or endorsed by them.
-
Their wordsEven more recently, we’ve begun to use mechanistic interpretability techniques to improve our safeguards and to conduct “audits” of new models before we release them, looking for evidence of deception, scheming, power-seeking, or a propensity to behave differently when being evaluated.
↗The Adolescence of Technology (January 2026)darioamodei.com
- 9 months earlier
-
Their wordsThese systems will be absolutely central to the economy, technology, and national security, and will be capable of so much autonomy that I consider it basically unacceptable for humanity to be totally ignorant of how they work.
↗The Urgency of Interpretability (April 2025)darioamodei.com