korrents

What Ilya Sutskever thinks about reinforcement learning

@ilya-sutskever · 28 positions · 0 changes of mind

Co-founder and chief scientist of Safe Superintelligence Inc., and previously co-founder and chief scientist of OpenAI.

Ilya Sutskever did not write this page.

We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.

7 dated positions, 2025, in their own words. Our reading of what Ilya Sutskever has said — not written or endorsed by them.

7 positions so far — this page is not yet offered to search engines.

  1. And one of the one thing you could do, and I think that's something that is done inadvertently, is that people take inspiration from the evals.

    Ilya Sutskever – We're moving from the age of scaling to the age of researchyoutube.com 1st of 28 in this recording

  2. I think for example our intuitive feeling of hunger is not succeeding in guiding us correctly in this world with an abundance of food.

    Ilya Sutskever – We're moving from the age of scaling to the age of researchyoutube.com 4th of 28 in this recording

  3. I want to like emphasize that I think the value function is something like it's going to make RL more efficient and I think that makes a difference but I think that anything you can do with a value function you can do without just more slowly.

    Ilya Sutskever – We're moving from the age of scaling to the age of researchyoutube.com 7th of 28 in this recording

  4. And so because scaling sucked out all the air in the room, everyone started to do the same thing. We got to the point where uh we are in a world where there are more companies than ideas by quite a bit.

    Ilya Sutskever – We're moving from the age of scaling to the age of researchyoutube.com 12th of 28 in this recording

  5. So there definitely for for research you need like definitely some amount of compute but it's far from obvious that you need the absolutely largest amount of compute ever for research.

    Ilya Sutskever – We're moving from the age of scaling to the age of researchyoutube.com 13th of 28 in this recording

  6. So the reason there has been no diversity I believe is because of pre-training. All the pre-trained models are the same pretty much because the pre-train on the same data.

    Ilya Sutskever – We're moving from the age of scaling to the age of researchyoutube.com 26th of 28 in this recording

    scaling laws

  7. Now the the thing is that selfplay at least the way it was done in the past when you have agents which are somehow compete with each other it's only good for developing a certain set of skills it is too narrow.

    Ilya Sutskever – We're moving from the age of scaling to the age of researchyoutube.com 27th of 28 in this recording