korrents

What Richard Sutton thinks about reinforcement learning

@richard-sutton · 25 positions · 0 changes of mind

Computer scientist who founded the field of reinforcement learning, co-author of the standard textbook, and winner of the 2024 Turing Award.

Richard Sutton did not write this page.

We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.

6 dated positions, 2025, in their own words. Our reading of what Richard Sutton has said — not written or endorsed by them.

6 positions so far — this page is not yet offered to search engines.

  1. reinforcement learning is about understanding your world whereas large language models are about mimicking people doing what people say you should do. They're not about figuring out what to do.

    Richard Sutton – Father of RL thinks LLMs are a dead endyoutube.com 2nd of 25 in this recording

    LLMs

  2. they have the ability to predict what a person would say they don't have the ability to predict what will happen

    Richard Sutton – Father of RL thinks LLMs are a dead endyoutube.com 3rd of 25 in this recording

    LLMs

  3. squirrels don't go to school. Squirrels can learn all about the world. It's absolutely obvious I would say that um supervised learning doesn't happen in animals.

    Richard Sutton – Father of RL thinks LLMs are a dead endyoutube.com 10th of 25 in this recording

  4. You can't have one child learn grow up and and learn about the world and then and then every new child has to repeat that process. Whereas with AIS, with a digital intelligence, you could hope to do it once and then copy it into the next one as a starting place.

    Richard Sutton – Father of RL thinks LLMs are a dead endyoutube.com 12th of 25 in this recording

  5. when you learn to play chess you have the grand the long-term goal is winning the game and yet you you can't you um you want to be able to learn from shorter term things like you know taking the your opponent's pieces um and so you do that by having a value function which predicts the long-term outcome

    Richard Sutton – Father of RL thinks LLMs are a dead endyoutube.com 13th of 25 in this recording

  6. Well, there's nothing in them which will cause it to generalize. Well, the gradient descent will cause them to find a solution to the problems they've seen. And if there's only one way to solve them, you know, they they'll do it. But there are many ways to solve it. Some which generalize well, some which generalize poorly. There's nothing in them in the algorithms that will cause them to generalize well.

    Richard Sutton – Father of RL thinks LLMs are a dead endyoutube.com 17th of 25 in this recording

    optimizers