A korrentour readingWhat is a korrent?
Reinforcement learning leaves models sharply jagged: on the rails of a verifiable domain they are superintelligent, off them everything meanders.
Drawn from what Andrej Karpathy said
What this subject means
reinforcement learning Training by reward: environments, verifiable tasks, value functions, and whether it works or merely beats what came before.