What Jan Leike thinks about AI alignment
Alignment researcher. Co-led OpenAI's superalignment team until 2024 and has led alignment work at Anthropic since; writes Musings on the Alignment Problem.
Jan Leike did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
5 dated positions, 2022 to 2026, in their own words. Our reading of what Jan Leike has said — not written or endorsed by them.
-
Their wordsThis is the hard problem of alignment we need to solve in order to succeed at building superintelligence, and to this day it is an unsolved problem.
↗Alignment is not solvedaligned.substack.com 2nd of 5 in this piece
-
Their wordsBut the goal we need to achieve is so much easier: we just need to build a model that’s as good as us at alignment research, and that we trust more than ourselves to do this research well because it’s sufficiently aligned.
↗Alignment is not solvedaligned.substack.com 3rd of 5 in this piece
- 12 months earlier
-
Their wordsMore generally, we should actually solve alignment instead of just trying to control misaligned AI.
↗Should we control AI instead of aligning it?aligned.substack.com 2nd of 2 in this piece
- 16 months earlier
-
Their wordsIf a model was capable of self-exfiltration, it would have the option to remove itself from your control.
↗Self-exfiltration is a key dangerous capabilityaligned.substack.com 2nd of 2 in this piece
- 18 months earlier
-
Their wordsThis is because the language modeling objective used for many recent large LMs-predicting the next token on a webpage from the internet-is different from the objective "follow the user's instructions helpfully and safely" (Radford et al.,, 2019; Brown et al.,, 2020; Fedus et al.,, 2021; Rae et al.,, 2021; Thoppilan et al.,, 2022). Thus, we say that the language modeling objective is misaligned.
↗Training language models to follow instructions with human feedback (with 19 co-authors)arxiv.org 4th of 4 in this piece