A korrentour readingWhat is a korrent?
Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores.
Drawn from what Sebastian Ruder said
What this subject means
scaling laws Whether more compute keeps buying more capability, and where the curve now bends -- pre-training, post-training, test-time.