korrents

A korrentour readingWhat is a korrent?

Very capable AI will be harder to align than current systems, because the loop of spotting a bad behaviour and patching the training that caused it breaks down.

Drawn from what Ryan Greenblatt said

A private bookmark. Not a position, and never counted.

What Ryan Greenblatt actually said

Word for word, with the source under each one. They did not write this page.

  1. Ryan Greenblatt

    Chief scientist at Redwood Research

    when AIs are extremely extremely capable my view is that those AIs will be harder to align than current systems. So for current systems, we have this feedback loop where we basically like we create an AI. We do some evaluations on it. We see that it has some kind of messed up behavior that we can kind of quickly understand. Then we like can like go look in training and be like, "Oh, the these training environments led to this problematic behavior. Let's like tweak that training data. Let's introduce some additional training data to like correct this other issue and then move forward from there." But in a regime where the AIs are extremely situationally aware, very very very very capable and um you know uh we don't necessarily understand what they're doing, this feedback loop breaks down.

Added to korrents 11 Aug 2026 · How quotes work · Something wrong? Tell us

Do you hold this korrent?Do you also believe this?

Sign in to record that you hold this, with a confidence number of your own.

Related korrents

Our reading — they may agree, disagree or merely touch the same thing. Closest first: a shared subject counts for most, then how near the wording is.

See this on the map →