korrents

korrents · piece

The core challenge of AI alignment is “steerability”

Helen Toner · 3 Apr 2025 · helentoner.substack.com

2 korrents from this piece

Helen Toner did not write this page.

Every claim below was made in this piece, quoted word for word and numbered in the order the piece makes them, so you can read it there rather than take our word for it. The sentence above each quote is our reading of the claim, not their wording. Each quote was checked against a stored copy of the page at build time; where the two differ, the quote is the fact.

  1. I see this as a totally fair question that totally misses the point of what “alignment” was trying to refer to: whether we’d be able to reliably steer advanced systems towards anything at all.
  2. But no one has (yet) been able to develop a model that is resistant to targeted “jailbreaking” that evades these restrictions.