korrents

Jan Leike

@jan-leike · 12 positions · 0 changes of mind

Alignment researcher. Co-led OpenAI's superalignment team until 2024 and has led alignment work at Anthropic since; writes Musings on the Alignment Problem.

Jan Leike did not write this page.

We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.

  1. But the most important lesson is that simple interventions are very effective at steering the model towards more aligned behavior.

    Alignment is not solvedaligned.substack.com 1st of 5 in this piece

  2. This is the hard problem of alignment we need to solve in order to succeed at building superintelligence, and to this day it is an unsolved problem.

    Alignment is not solvedaligned.substack.com 2nd of 5 in this piece

    AI alignment

  3. But the goal we need to achieve is so much easier: we just need to build a model that’s as good as us at alignment research, and that we trust more than ourselves to do this research well because it’s sufficiently aligned.

    Alignment is not solvedaligned.substack.com 3rd of 5 in this piece

    AI alignment

  4. We are starting to automate AI research and the recursive self-improvement process has begun.

    Alignment is not solvedaligned.substack.com 4th of 5 in this piece

    recursive self-improvement

  5. In fact, making an evil version of Claude that’s just as smart and agentic would be pretty easy.

    Alignment is not solvedaligned.substack.com 5th of 5 in this piece

    Anthropic

  6. 12 months earlier
  7. However, without having a good handle on elicitation, we cannot be confident that our control techniques are effective.

    Should we control AI instead of aligning it?aligned.substack.com 1st of 2 in this piece

  8. More generally, we should actually solve alignment instead of just trying to control misaligned AI.

    Should we control AI instead of aligning it?aligned.substack.com 2nd of 2 in this piece

    AI alignment

  9. 3 months earlier
  10. This is especially severe for open weight models, where deployment is essentially irreversible and mitigations can’t be improved later.

    Two alignment threat modelsaligned.substack.com 1st of 3 in this piece

  11. Therefore it’s simply intractable to meaningfully oversee all but a tiny fraction of our models’ behavior with actual humans.

    Two alignment threat modelsaligned.substack.com 2nd of 3 in this piece

  12. In other words, current models are probably severely under-elicited.

    Two alignment threat modelsaligned.substack.com 3rd of 3 in this piece

  13. 14 months earlier
  14. Moreover, model exfiltration is likely impossible to reverse.

    Self-exfiltration is a key dangerous capabilityaligned.substack.com 1st of 2 in this piece

  15. If a model was capable of self-exfiltration, it would have the option to remove itself from your control.

    Self-exfiltration is a key dangerous capabilityaligned.substack.com 2nd of 2 in this piece

    AI alignment