korrents

A korrentour readingWhat is a korrent?

Learning from a reward ten years away is a solved problem: a value function trained by temporal-difference learning rewards the steps along the way.

Drawn from what Richard Sutton said

A private bookmark. Not a position, and never counted.

What Richard Sutton actually said

Word for word, with the source under each one. They did not write this page.

  1. Richard Sutton

    Computer scientist who founded the field of reinforcement learning

    when you learn to play chess you have the grand the long-term goal is winning the game and yet you you can't you um you want to be able to learn from shorter term things like you know taking the your opponent's pieces um and so you do that by having a value function which predicts the long-term outcome

Added to korrents 26 Sept 2025 · How quotes work · Something wrong? Tell us

Do you hold this korrent?Do you also believe this?

Sign in to record that you hold this, with a confidence number of your own.

Related korrents

Our reading — they may agree, disagree or merely touch the same thing. Closest first: a shared subject counts for most, then how near the wording is.

On the map

Loading the map… or open it on its own page

Open the map on its own page →