A korrentour readingWhat is a korrent?
RL environments exist to make a model generalise, not to teach it each skill one at a time — exactly as pre-training does.
Drawn from what Dario Amodei said
scaling laws Whether more compute keeps buying more capability, and where the curve now bends -- pre-training, post-training, test-time.
reinforcement learning Training by reward: environments, verifiable tasks, value functions, and whether it works or merely beats what came before.