korrents

korrents · piece

The case for ensuring that powerful AIs are controlled

Buck Shlegeris · 7 May 2024 · redwoodresearch.substack.com

8 korrents from this piece

Buck Shlegeris did not write this page.

Every claim below was made in this piece, quoted word for word and numbered in the order the piece makes them, so you can read it there rather than take our word for it. The sentence above each quote is our reading of the claim, not their wording. Each quote was checked against a stored copy of the page at build time; where the two differ, the quote is the fact.

  1. That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures.
  2. We think no fundamental research breakthroughs are required for labs to implement safety measures that meet our standard for AI control for early transformatively useful AIs; we think that meeting our standard would substantially reduce the risks posed by intentional subversion.
  3. So when evaluating control, we should count catching an AI red-handed as a win condition.
  4. Because evaluating control just requires evaluating capabilities, it's far easier to robustly evaluate than alignment.
  5. We're advocating that companies handle risk from scheming models in a similar way–striving to ensure that they'll be safe even if their alignment efforts fail to prevent models from scheming.
  6. In particular, suppose that our control evaluations directly estimate a less than 1% chance of catastrophe if our untrusted AIs are scheming, then the actual risk conditional on scheming is probably more like 5% to 20% due to the potential for failures in the evaluation.
  7. The basic problem with evaluating alignment is that no matter what behaviors you observe, you have to worry that your model is just acting that way in order to make you think that it is aligned.
  8. AI control (with only black-box techniques) seems like a fundamentally limited approach.