A korrentour readingWhat is a korrent?
Long-context evaluation benchmarks are likely contaminated by model training data, making them unreliable as a sole evaluation method
Drawn from what Eugene Yan said
Subjectbenchmarks
A korrentour readingWhat is a korrent?
Drawn from what Eugene Yan said
Word for word, with the source under each one. They did not write this page.
Member of technical staff at Anthropic
Since these datasets are likely already part of model training data, we shouldn't rely solely on them to evaluate our Q&A system.
Where this was said
↗Evaluating Long-Context Question & Answer Systemseugeneyan.com
All 3 korrents from this piecethis one is 3rd
Added to korrents 22 Jun 2025 · How quotes work · Something wrong? Tell us
Our reading — they may agree, disagree or merely touch the same thing. Closest first: a shared subject counts for most, then how near the wording is.