korrents

Buck Shlegeris

@buck-shlegeris · 8 positions · 0 changes of mind

CEO of Redwood Research, where he works on AI control: making sure a powerful model cannot cause a catastrophe even if it is misaligned and trying to.

Buck Shlegeris did not write this page.

We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.

  1. That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures.

    The case for ensuring that powerful AIs are controlledredwoodresearch.substack.com 1st of 8 in this piece

    AI alignment

  2. We think no fundamental research breakthroughs are required for labs to implement safety measures that meet our standard for AI control for early transformatively useful AIs; we think that meeting our standard would substantially reduce the risks posed by intentional subversion.

    The case for ensuring that powerful AIs are controlledredwoodresearch.substack.com 2nd of 8 in this piece

  3. So when evaluating control, we should count catching an AI red-handed as a win condition.

    The case for ensuring that powerful AIs are controlledredwoodresearch.substack.com 3rd of 8 in this piece

  4. Because evaluating control just requires evaluating capabilities, it's far easier to robustly evaluate than alignment.

    The case for ensuring that powerful AIs are controlledredwoodresearch.substack.com 4th of 8 in this piece

    AI alignment

  5. We're advocating that companies handle risk from scheming models in a similar way–striving to ensure that they'll be safe even if their alignment efforts fail to prevent models from scheming.

    The case for ensuring that powerful AIs are controlledredwoodresearch.substack.com 5th of 8 in this piece

    AI alignment

  6. In particular, suppose that our control evaluations directly estimate a less than 1% chance of catastrophe if our untrusted AIs are scheming, then the actual risk conditional on scheming is probably more like 5% to 20% due to the potential for failures in the evaluation.

    The case for ensuring that powerful AIs are controlledredwoodresearch.substack.com 6th of 8 in this piece

  7. The basic problem with evaluating alignment is that no matter what behaviors you observe, you have to worry that your model is just acting that way in order to make you think that it is aligned.

    The case for ensuring that powerful AIs are controlledredwoodresearch.substack.com 7th of 8 in this piece

    AI alignment

  8. AI control (with only black-box techniques) seems like a fundamentally limited approach.

    The case for ensuring that powerful AIs are controlledredwoodresearch.substack.com 8th of 8 in this piece

    AI alignment