korrents

korrents · piece

Alignment is not solved

Jan Leike · 22 Jan 2026 · aligned.substack.com

5 korrents from this piece

Jan Leike did not write this page.

Every claim below was made in this piece, quoted word for word and numbered in the order the piece makes them, so you can read it there rather than take our word for it. The sentence above each quote is our reading of the claim, not their wording. Each quote was checked against a stored copy of the page at build time; where the two differ, the quote is the fact.

  1. But the most important lesson is that simple interventions are very effective at steering the model towards more aligned behavior.
  2. This is the hard problem of alignment we need to solve in order to succeed at building superintelligence, and to this day it is an unsolved problem.
  3. But the goal we need to achieve is so much easier: we just need to build a model that’s as good as us at alignment research, and that we trust more than ourselves to do this research well because it’s sufficiently aligned.
  4. We are starting to automate AI research and the recursive self-improvement process has begun.
  5. In fact, making an evil version of Claude that’s just as smart and agentic would be pretty easy.