What Helen Toner thinks about AI alignment
Director of strategy at Georgetown's Center for Security and Emerging Technology and a former member of OpenAI's board. Writes Rising Tide, on how AI actually progresses and what policy can do about it.
Helen Toner did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
7 dated positions, 2025, in their own words. Our reading of what Helen Toner has said — not written or endorsed by them.
-
Their wordsWe should think of the gap between frontier models and proliferated models as an adaptation buffer: a limited time window when we know what bad actors will soon be able to use AI for, which gives us a chance to implement defensive measures that increase society’s resilience to the danger.
↗Nonproliferation is the wrong approach to AI misusehelentoner.substack.com 2nd of 5 in this piece
-
Their wordsBut we do not have experience with the kinds of problems we’ll face if we build machines that are smarter and more capable than us, so I feel much less sanguine that an adaptation-based strategy will help us there.
↗Nonproliferation is the wrong approach to AI misusehelentoner.substack.com 4th of 5 in this piece
-
Their wordsBy trying to use misuse as a fig leaf for their real concerns, they end up sounding much less credible than if they just tried to argue for what they actually meant.
↗Nonproliferation is the wrong approach to AI misusehelentoner.substack.com 5th of 5 in this piece
- 2 days earlier
-
Their wordsI see this as a totally fair question that totally misses the point of what “alignment” was trying to refer to: whether we’d be able to reliably steer advanced systems towards anything at all.
↗The core challenge of AI alignment is “steerability”helentoner.substack.com 1st of 2 in this piece
-
Their wordsBut no one has (yet) been able to develop a model that is resistant to targeted “jailbreaking” that evades these restrictions.
↗The core challenge of AI alignment is “steerability”helentoner.substack.com 2nd of 2 in this piece
- 2 days earlier
-
Their wordsDismissing discussion of AGI, human-level AI, transformative AI, superintelligence, etc. as “science fiction” should be seen as a sign of total unseriousness.
↗"Long" timelines to advanced AI have gotten crazy shorthelentoner.substack.com 1st of 3 in this piece
-
Their wordsWe need to leap into action on many of the same things that could help if it does turn out that we only have a few years.
↗"Long" timelines to advanced AI have gotten crazy shorthelentoner.substack.com 3rd of 3 in this piece