korrents

korrents · piece

Two alignment threat models

Jan Leike · 8 Nov 2024 · aligned.substack.com

3 korrents from this piece

Jan Leike did not write this page.

Every claim below was made in this piece, quoted word for word and numbered in the order the piece makes them, so you can read it there rather than take our word for it. The sentence above each quote is our reading of the claim, not their wording. Each quote was checked against a stored copy of the page at build time; where the two differ, the quote is the fact.

  1. This is especially severe for open weight models, where deployment is essentially irreversible and mitigations can’t be improved later.
  2. Therefore it’s simply intractable to meaningfully oversee all but a tiny fraction of our models’ behavior with actual humans.
  3. In other words, current models are probably severely under-elicited.