korrents

← People and their mental models

Buck Shlegeris's mental models

2 claims Buck Shlegeris made fit 1 mental model. Most often: Margin of safety. Everything they said here.

Models we see in what they say

Our reading: their claim applies the idea without naming it. The claim is theirs; filing it here is ours.

Margin of safety

Leave room for being wrong, because sometimes you will be.

Used by 25 others

A control evaluation that reports under one per cent risk should be read as several per cent, because the evaluation can itself fail.

  1. Buck Shlegeris CEO of Redwood Research In particular, suppose that our control evaluations directly estimate a less than 1% chance of catastrophe if our untrusted AIs are scheming, then the actual risk conditional on scheming is probably more like 5% to 20% due to the potential for failures in the evaluation. The case for ensuring that powerful AIs are controlledredwoodresearch.substack.com · 7 May 2024All korrents from this piece
    In particular, suppose that our control evaluations directly estimate a less than 1% chance of catastrophe if our untrusted AIs are scheming, then the actual risk conditional on scheming is probably more like 5% to 20% due to the potential for failures in the evaluation.

Risk from a scheming model should be managed the way insider security risk is: assume some of them are compromised and stay safe anyway.

  1. Buck Shlegeris CEO of Redwood Research We're advocating that companies handle risk from scheming models in a similar way–striving to ensure that they'll be safe even if their alignment efforts fail to prevent models from scheming. The case for ensuring that powerful AIs are controlledredwoodresearch.substack.com · 7 May 2024All korrents from this piece
    We're advocating that companies handle risk from scheming models in a similar way–striving to ensure that they'll be safe even if their alignment efforts fail to prevent models from scheming.