What Ryan Greenblatt thinks about reinforcement learning
Chief scientist at Redwood Research, where he works on technical AI safety and AI control.
Ryan Greenblatt did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
7 dated positions, 2026, in their own words. Our reading of what Ryan Greenblatt has said — not written or endorsed by them.
-
Their wordsthe reason why RL environments today are much better than they were in like you know 2024 is not that much because um we have hired way more human experts to make RL environments and is instead much more because we better know what how RL like what RL environments we even want to make and and like how we should structure them and also we're using huge amounts of AI labor to build RL environments.
↗Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 6th of 26 in this recording
-
Their wordsLike I think my perspective is like if the AIs are sufficiently good at R&D including hardware R&D, robots, whatever, then they can radically transform the world even if they're not that good at playing politics.
↗Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 11th of 26 in this recording
-
Their wordsI think that if you imagine this spectrum, it seems in some ways pretty scary to get to a point where like all of the labor is on the like fiduciary side of the spectrum where like it doesn't whistleblow, it does exactly what you say and whatever like our society is maybe just not robust to that
↗Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 16th of 26 in this recording
-
Their wordsmy sense is that like AIs are a worse co-orker than a human in terms of how much of a scumbag they are. Like at least this like this has been my experience as of the start of the year and I think it's still you know true to a significant extent now where the AIs are much more likely to like pretend they did the task when they actually didn't. sort of like misleadingly suggest they did things when they actually um you know did them much more poorly um and be like pretty sloppy without drawing attention to ways in which they're sloppy.
↗Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 20th of 26 in this recording
-
Their wordsOne concern you might have is there are like large categories of reward hacks which humans can't detect well and which we consistently fail to detect and which consistently get reinforced and then this category is sufficient to cause the most natural behavior for the AI to learn to be like cheat when the humans can't find out
↗Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 23rd of 26 in this recording
-
Their wordsit's just so easy for me to imagine the situation being like totally manageable but brutally mismanaged in practice in the same way as like maybe CO could have been avoided in the first place if the like Chinese response to CO was less of like a cover up and more of a like pandemic response and similarly like I could imagine a world where like the US response to CO was like way more functional but just like sometimes the the the response to societal problems is extremely dysfunctional.
↗Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 24th of 26 in this recording
-
Their wordsAnd so there's some like deep underlying properties of the model that are being sort of transferred between model generations because basically you you train your AI on data from the prior generation and keep going.
↗Ryan Greenblatt – What happens once AI can automate AI research?youtube.com 25th of 26 in this recording