A korrentour readingWhat is a korrent?
Reinforcement learning from verifiable rewards does not reliably carry a model's pretrained ethical understanding into its actual behavior.
Drawn from what Samuel Hammond said
What these subjects mean
scaling laws Whether more compute keeps buying more capability, and where the curve now bends -- pre-training, post-training, test-time.
reinforcement learning Training by reward: environments, verifiable tasks, value functions, and whether it works or merely beats what came before.