A korrentour readingWhat is a korrent?
RLHF is not a tax on capability: preference tuning also raises maths and code scores, which is why the labs keep reaching for it.
Drawn from what Nathan Lambert said
What this subject means
reinforcement learning Training by reward: environments, verifiable tasks, value functions, and whether it works or merely beats what came before.