korrents

A korrentour readingWhat is a korrent?

RL environments exist to make a model generalise, not to teach it each skill one at a time — exactly as pre-training does.

Drawn from what Dario Amodei said

scaling laws Whether more compute keeps buying more capability, and where the curve now bends -- pre-training, post-training, test-time.

reinforcement learning Training by reward: environments, verifiable tasks, value functions, and whether it works or merely beats what came before.

A private bookmark. Not a position, and never counted.

What Dario Amodei actually said

Word for word, with the source under each one. They did not write this page.

  1. Dario Amodei

    Co-founder and CEO of Anthropic

    the goal is not to teach the model every possible skill within RL just as we don't do that within pre-training, right? Within pre-training, we're not trying to expose the model to, you know, every every possible you know, way that words could be put together, right? You know, we're it's it's rather that the model trains on a lot of things and then and then it reaches generalization across pre-training, right?

Added to korrents 13 Feb 2026 · How quotes work · Something wrong? Tell us

Do you hold this korrent?Do you also believe this?

Sign in to record that you hold this, with a confidence number of your own.

Related korrents

Our reading — they may agree, disagree or merely touch the same thing. Closest first: a shared subject counts for most, then how near the wording is.

See this on the map →