A korrentour readingWhat is a korrent?
Today's RL environments are better than 2024's because we learned what to build and put AI labour on building it, not because labs hired more human experts.
Drawn from what Ryan Greenblatt said
What this subject means
reinforcement learning Training by reward: environments, verifiable tasks, value functions, and whether it works or merely beats what came before.