korrents

korrents · Dwarkesh Podcast

Ryan Greenblatt – What happens once AI can automate AI research?

Ryan Greenblatt · 2h 12m · youtube.com

26 korrents from this recording

1h2h
Ryan Greenblatt did not write this page.

Every claim below is a statement made in this recording, quoted word for word and linked to the second it was said, so you can hear it rather than take our word for it. The wording comes from the transcript published alongside the recording; the sentence above each quote is our reading of the claim, not their wording.

  1. 0:01:14 · watch on youtube.com

    Maybe my sort of median expectation is something like uh four or five years of AI progress in a single year.
  2. 2 min later
  3. 0:03:09 · watch on youtube.com

    I would say that I expect like full automation of AR&D perhaps somewhere around like 2031 2030 and then getting to like the like beats all humans on the job milestone. Maybe I expect median around 2033
  4. 8 min later
  5. 0:11:35 · watch on youtube.com

    Second, I think ML is a very shallow domain relative to math. So I think in math there's much more of a you find some true deep abstraction um and then like that like if you really understand that thing which is hard to understand then you get somewhere
  6. 8 min later
  7. 0:19:05 · watch on youtube.com

    Basically, the story would end up being that to get five years of AI progress, you're probably going to need around I would say like maybe eight years of algorithmic progress very roughly. Um, which is a lot a lot of algorithmic progress.
  8. 1 min later
  9. 0:20:18 · watch on youtube.com

    So my sense is that scaling up the amount of effort spent on getting expert human data has not been hugely important for AI R&D in general.
  10. 1 min later
  11. 0:21:17 · watch on youtube.com

    the reason why RL environments today are much better than they were in like you know 2024 is not that much because um we have hired way more human experts to make RL environments and is instead much more because we better know what how RL like what RL environments we even want to make and and like how we should structure them and also we're using huge amounts of AI labor to build RL environments.
  12. 6 min later
  13. 0:27:16 · watch on youtube.com

    I think most domains are fundamentally pretty shallow where like a very smart generalist who's good at like a a limited subset of core skills can like get going pretty quickly.
  14. 7 min later
  15. 0:34:10 · watch on youtube.com

    Like the thing that I think is most likely to be sort of the bottleneck in terms of like the AI are really good at verifiable domains but not not at doing the actual thing is just like big experiments. You only get a few tries um well a few is maybe a bit understated but like basically like historically R&D has been driven by doing near frontier scale experiments and that has been pretty important and like actually doing the one big training run where you decide exactly what to include in that.
  16. 1 min later
  17. 0:34:53 · watch on youtube.com

    I think one reason why um the the AIs have been scaled up less than you would have otherwise expected and like for example cost of of per token hasn't increased as much as you might have thought is because there's a benefit to doing more of your um work at small scale where you can run more training runs and get more cycles in
  18. 5 min later
  19. 0:39:40 · watch on youtube.com

    it's really hard for me to think of examples of cognitive tasks humans do where we're not seeing some transfer from AI improving.
  20. 7 min later
  21. 0:46:22 · watch on youtube.com

    Like I think my perspective is like if the AIs are sufficiently good at R&D including hardware R&D, robots, whatever, then they can radically transform the world even if they're not that good at playing politics.
  22. 7 min later
  23. 0:53:06 · watch on youtube.com

    I do I do think that I wish that sort of my preferred constitution or like the way I would orient towards this like the thing I would prefer would be more like Claude is like look it would be structurally good for the way this technology work like the constitution should be like it would be structurally good for the way this technology works to be that AIS are like good fiduciaries, good representatives, the equivalent of a lawyer for a user
  24. 1 min later
  25. 0:54:32 · watch on youtube.com

    we are making a trade-off where because we don't have very good alignment technology. We are going to like make an alien mind with its own values and then gamble on that to some extent rather than doing this other approach of making like a tool that pursues individual user intention.
  26. 5 min later
  27. 0:59:21 · watch on youtube.com

    I think this constitution is in some sense very compatible with Claude doing huge amounts of power seeking because it thinks that will result in better outcomes.
  28. 0:59:50 · watch on youtube.com

    But it's not very hard to imagine a situation in which the sort of long run values sink in deeper than the prohibitions against takeover.
  29. 7 min later
  30. 1:06:26 · watch on youtube.com

    I think that if you imagine this spectrum, it seems in some ways pretty scary to get to a point where like all of the labor is on the like fiduciary side of the spectrum where like it doesn't whistleblow, it does exactly what you say and whatever like our society is maybe just not robust to that
  31. 2 min later
  32. 1:08:01 · watch on youtube.com

    the most powerful actors for whom this is the biggest concern if these guard rails or the constitution or whatever is getting in the way that will just get steamrolled and so the constitution will only be you know hitting the everyday man rather than hitting governments.
  33. 5 min later
  34. 1:12:38 · watch on youtube.com

    when AIs are extremely extremely capable my view is that those AIs will be harder to align than current systems. So for current systems, we have this feedback loop where we basically like we create an AI. We do some evaluations on it. We see that it has some kind of messed up behavior that we can kind of quickly understand. Then we like can like go look in training and be like, "Oh, the these training environments led to this problematic behavior. Let's like tweak that training data. Let's introduce some additional training data to like correct this other issue and then move forward from there." But in a regime where the AIs are extremely situationally aware, very very very very capable and um you know uh we don't necessarily understand what they're doing, this feedback loop breaks down.
  35. 14 min later
  36. 1:26:14 · watch on youtube.com

    my expectation is what we would see from then is that the rate of problematic behavior would decrease uh and would just keep decreasing and decrease at a pretty fast rate while simultaneously the worst things that the AIS would sometimes do would get more extreme, more egregious, and more scary.
  37. 3 min later
  38. 1:29:02 · watch on youtube.com

    my sense is that like AIs are a worse co-orker than a human in terms of how much of a scumbag they are. Like at least this like this has been my experience as of the start of the year and I think it's still you know true to a significant extent now where the AIs are much more likely to like pretend they did the task when they actually didn't. sort of like misleadingly suggest they did things when they actually um you know did them much more poorly um and be like pretty sloppy without drawing attention to ways in which they're sloppy.
  39. 4 min later
  40. 1:33:04 · watch on youtube.com

    I would also note that my sense is that like the place where the misalignment most lives is the place where you're trying to really push the eyes hard and get them to like do work that's really on the cutting edge of what they are capable of
  41. 5 min later
  42. 1:38:02 · watch on youtube.com

    And so basically everything that we can verify reasonably well with some feedback loop, the AIS are doing pretty well on. And that's sufficient to make AR and D go quite fast and to continue. But there's some parts of of developing uh aligned and safe AIs that are more subtle, hard to check, depend on, you know, detailed in the weeds things.
  43. 3 min later
  44. 1:40:37 · watch on youtube.com

    One concern you might have is there are like large categories of reward hacks which humans can't detect well and which we consistently fail to detect and which consistently get reinforced and then this category is sufficient to cause the most natural behavior for the AI to learn to be like cheat when the humans can't find out
  45. 23 min later
  46. 2:03:29 · watch on youtube.com

    it's just so easy for me to imagine the situation being like totally manageable but brutally mismanaged in practice in the same way as like maybe CO could have been avoided in the first place if the like Chinese response to CO was less of like a cover up and more of a like pandemic response and similarly like I could imagine a world where like the US response to CO was like way more functional but just like sometimes the the the response to societal problems is extremely dysfunctional.
  47. 3 min later
  48. 2:06:21 · watch on youtube.com

    And so there's some like deep underlying properties of the model that are being sort of transferred between model generations because basically you you train your AI on data from the prior generation and keep going.
  49. 2 min later
  50. 2:08:04 · watch on youtube.com

    By 2040 um let's see uh maybe around 35 or 40%.