korrents

korrents · Dwarkesh Podcast

Richard Sutton – Father of RL thinks LLMs are a dead end

Richard Sutton · 1h 07m · youtube.com

25 korrents from this recording

1h
Richard Sutton did not write this page.

Every claim below is a statement made in this recording, quoted word for word and linked to the second it was said, so you can hear it rather than take our word for it. The wording comes from the transcript published alongside the recording; the sentence above each quote is our reading of the claim, not their wording.

  1. 0:00:00 · watch on youtube.com

    Why are you trying to distinguish humans? Humans are animals. What we have in common is more interesting. What distinguishes us, we should be paying less attention to.
  2. 2 min later
  3. 0:01:56 · watch on youtube.com

    reinforcement learning is about understanding your world whereas large language models are about mimicking people doing what people say you should do. They're not about figuring out what to do.
  4. 1 min later
  5. 0:02:53 · watch on youtube.com

    they have the ability to predict what a person would say they don't have the ability to predict what will happen
  6. 2 min later
  7. 0:05:09 · watch on youtube.com

    So there's no ground truth. You can't have prior knowledge if you don't have ground truth because the prior knowledge is supposed to be a hint or an initial belief about what the truth is.
  8. 3 min later
  9. 0:07:47 · watch on youtube.com

    For me, having a goal is the essence of intelligence
  10. 5 min later
  11. 0:13:00 · watch on youtube.com

    But in fact and in practice it has always turned out to be bad because people get locked into the human knowledge approach and they psychologically or you know now I'm now I'm speculating why it is but this is what has always happened.
  12. 0:13:23 · watch on youtube.com

    The scalable method is you learn from experience. Um you uh you you try things, you see what you see what works. No one no one has to tell you. First of all, you have a goal. So without a goal, uh there's no sense of right or wrong or better or worse. So large language models are trying to get by without having a goal or a sense of better or worse. That's just, you know, it's exactly starting in the wrong place.
  13. 4 min later
  14. 0:17:07 · watch on youtube.com

    So I don't think uh learning is really about training. I think learning is about about learning. It's about an active process. The child tries things and sees what happens.
  15. 0:17:27 · watch on youtube.com

    If you go to look about how psychologists think about learning, there's nothing like uh imitation. Maybe there are some extreme cases where humans might do that or appear to do that, but there's no basic animal learning process called imitation.
  16. 1 min later
  17. 0:18:47 · watch on youtube.com

    squirrels don't go to school. Squirrels can learn all about the world. It's absolutely obvious I would say that um supervised learning doesn't happen in animals.
  18. 1 min later
  19. 0:19:52 · watch on youtube.com

    if we understood a squirrel I think we'd have a we'd be almost all the way there to understanding human intelligence. The the language part is just a a small veneer on the surface.
  20. 8 min later
  21. 0:27:24 · watch on youtube.com

    You can't have one child learn grow up and and learn about the world and then and then every new child has to repeat that process. Whereas with AIS, with a digital intelligence, you could hope to do it once and then copy it into the next one as a starting place.
  22. 1 min later
  23. 0:28:39 · watch on youtube.com

    when you learn to play chess you have the grand the long-term goal is winning the game and yet you you can't you um you want to be able to learn from shorter term things like you know taking the your opponent's pieces um and so you do that by having a value function which predicts the long-term outcome
  24. 8 min later
  25. 0:36:45 · watch on youtube.com

    We don't have any methods that are good at that. What we have are people um try different things and they they settle on something that that uh a representation that that transfers well or they generalize as well. But we have no we don't have any automated techniques to promote. we have very few automated techniques to promote transfer and they're not none of them are used in in modern deep learning.
  26. 1 min later
  27. 0:37:47 · watch on youtube.com

    so we know deep learning is really bad at this for example we know that if you train on some new thing it will often catastrophically interfere with all the old things that you that you knew
  28. 2 min later
  29. 0:39:19 · watch on youtube.com

    We don't we don't really know what information they had prior. We are we have to guess because they've been fed so much. This is one reason why they're not a good way to do science. Uh it's just so uncontrolled, so unknown.
  30. 1 min later
  31. 0:40:41 · watch on youtube.com

    Well, there's nothing in them which will cause it to generalize. Well, the gradient descent will cause them to find a solution to the problems they've seen. And if there's only one way to solve them, you know, they they'll do it. But there are many ways to solve it. Some which generalize well, some which generalize poorly. There's nothing in them in the algorithms that will cause them to generalize well.
  32. 7 min later
  33. 0:47:17 · watch on youtube.com

    I I really view myself as a classicist rather than as a contrarian. I go to what what the larger community of of thinkers about the mind have always thought.
  34. 4 min later
  35. 0:50:50 · watch on youtube.com

    The bitter lesson. Oh, who cares about that? That's that's an empirical observation about a particular period in history. 70 years in history no longer doesn't necessarily have to apply the next 70 years.
  36. 2 min later
  37. 0:52:32 · watch on youtube.com

    But it will not be that easy, as easy as you're imagining because uh that you can lose your mind this way. If you you pull in something from the outside and build it into your into your inner thinking, uh, it could take over you. It could change you.
  38. 2 min later
  39. 0:54:48 · watch on youtube.com

    So I do think succession to digital or digital intelligence or augmented humans is inevitable.
  40. 3 min later
  41. 0:57:29 · watch on youtube.com

    then we're entering the age of design where because our AIs are designed our our our all of our physical objects are designed our buildings are designed our technology is designed and we're we're designing now uh AIs things that can be intelligent themselves and that are themselves capable of design
  42. 2 min later
  43. 0:59:17 · watch on youtube.com

    It's our choice whether we should say oh they are our offspring and we should be proud of them and we should celebrate their achievements or we should we could say oh no they're not us and we should be horrified.
  44. 2 min later
  45. 1:01:23 · watch on youtube.com

    A lot of it has to do with just how you feel about change. Um, and if you think the current situation is really really good, then you're uh more likely to be suspicious of change and averse to change than if you think um it's imperfect. And I think it's imperfect. In fact, I think it's pretty bad.
  46. 1 min later
  47. 1:02:40 · watch on youtube.com

    we al also though should recognize the limits, our limits. And we're I think we want to avoid the feeling of entitlement. Avoid the feeling, oh, we are here first. We should always have it in a good way.