A korrentour readingWhat is a korrent?
Scaling up GPT will never produce AGI, because a model trained on cross-entropy loss cannot get there; reinforcement learning in rich environments is required.
Drawn from what George Hotz said
What this subject means
reinforcement learning Training by reward: environments, verifiable tasks, value functions, and whether it works or merely beats what came before.