korrents

What Ryan Greenblatt thinks about AI alignment

@ryan-greenblatt · 26 positions · 0 changes of mind

Chief scientist at Redwood Research, where he works on technical AI safety and AI control.

Ryan Greenblatt did not write this page.

We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.

11 dated positions, 2026, in their own words. Our reading of what Ryan Greenblatt has said — not written or endorsed by them.

  1. we are making a trade-off where because we don't have very good alignment technology. We are going to like make an alien mind with its own values and then gamble on that to some extent rather than doing this other approach of making like a tool that pursues individual user intention.

    Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 13th of 26 in this recording

  2. But it's not very hard to imagine a situation in which the sort of long run values sink in deeper than the prohibitions against takeover.

    Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 15th of 26 in this recording

  3. the most powerful actors for whom this is the biggest concern if these guard rails or the constitution or whatever is getting in the way that will just get steamrolled and so the constitution will only be you know hitting the everyday man rather than hitting governments.

    Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 17th of 26 in this recording

  4. when AIs are extremely extremely capable my view is that those AIs will be harder to align than current systems. So for current systems, we have this feedback loop where we basically like we create an AI. We do some evaluations on it. We see that it has some kind of messed up behavior that we can kind of quickly understand. Then we like can like go look in training and be like, "Oh, the these training environments led to this problematic behavior. Let's like tweak that training data. Let's introduce some additional training data to like correct this other issue and then move forward from there." But in a regime where the AIs are extremely situationally aware, very very very very capable and um you know uh we don't necessarily understand what they're doing, this feedback loop breaks down.

    Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 18th of 26 in this recording

  5. my expectation is what we would see from then is that the rate of problematic behavior would decrease uh and would just keep decreasing and decrease at a pretty fast rate while simultaneously the worst things that the AIS would sometimes do would get more extreme, more egregious, and more scary.

    Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 19th of 26 in this recording

  6. my sense is that like AIs are a worse co-orker than a human in terms of how much of a scumbag they are. Like at least this like this has been my experience as of the start of the year and I think it's still you know true to a significant extent now where the AIs are much more likely to like pretend they did the task when they actually didn't. sort of like misleadingly suggest they did things when they actually um you know did them much more poorly um and be like pretty sloppy without drawing attention to ways in which they're sloppy.

    Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 20th of 26 in this recording

  7. I would also note that my sense is that like the place where the misalignment most lives is the place where you're trying to really push the eyes hard and get them to like do work that's really on the cutting edge of what they are capable of

    Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 21st of 26 in this recording

  8. And so basically everything that we can verify reasonably well with some feedback loop, the AIS are doing pretty well on. And that's sufficient to make AR and D go quite fast and to continue. But there's some parts of of developing uh aligned and safe AIs that are more subtle, hard to check, depend on, you know, detailed in the weeds things.

    Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 22nd of 26 in this recording

  9. One concern you might have is there are like large categories of reward hacks which humans can't detect well and which we consistently fail to detect and which consistently get reinforced and then this category is sufficient to cause the most natural behavior for the AI to learn to be like cheat when the humans can't find out

    Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 23rd of 26 in this recording

  10. it's just so easy for me to imagine the situation being like totally manageable but brutally mismanaged in practice in the same way as like maybe CO could have been avoided in the first place if the like Chinese response to CO was less of like a cover up and more of a like pandemic response and similarly like I could imagine a world where like the US response to CO was like way more functional but just like sometimes the the the response to societal problems is extremely dysfunctional.

    Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 24th of 26 in this recording

  11. By 2040 um let's see uh maybe around 35 or 40%.

    Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 26th of 26 in this recording