A korrentour readingWhat is a korrent?
When pre-training data is limited, it is better to train a smaller model for multiple epochs than to train a larger model on unique data once.
Drawn from what Sebastian Ruder said
What this subject means
scaling laws Whether more compute keeps buying more capability, and where the curve now bends -- pre-training, post-training, test-time.