A korrentour readingWhat is a korrent?
The knowledge a model soaks up in pre-training is holding it back; what we actually want is the intelligence with the knowledge stripped out.
Drawn from what Andrej Karpathy said
What this subject means
scaling laws Whether more compute keeps buying more capability, and where the curve now bends -- pre-training, post-training, test-time.