A korrentour readingWhat is a korrent?
Humans keep their place in AI as judges rather than authors, because telling which of two answers is better is far easier than writing a good one.
Drawn from what Nathan Lambert said
What this subject means
reinforcement learning Training by reward: environments, verifiable tasks, value functions, and whether it works or merely beats what came before.