A control evaluation that reports under one per cent risk should be read as several per cent, because the evaluation can itself fail.
Buck Shlegeris CEO of Redwood Research In particular, suppose that our control evaluations directly estimate a less than 1% chance of catastrophe if our untrusted AIs are scheming, then the actual risk conditional on scheming is probably more like 5% to 20% due to the potential for failures in the evaluation. The case for ensuring that powerful AIs are controlledredwoodresearch.substack.com · 7 May 2024All korrents from this piece
Their wordsIn particular, suppose that our control evaluations directly estimate a less than 1% chance of catastrophe if our untrusted AIs are scheming, then the actual risk conditional on scheming is probably more like 5% to 20% due to the potential for failures in the evaluation.