korrents

A korrentour readingWhat is a korrent?

Long-context evaluation benchmarks are likely contaminated by model training data, making them unreliable as a sole evaluation method

Drawn from what Eugene Yan said

benchmarks

What Eugene Yan actually said

Word for word, with the source under each one. They did not write this page.

  1. Eugene Yan

    Member of technical staff at Anthropic

    Since these datasets are likely already part of model training data, we shouldn't rely solely on them to evaluate our Q&A system.

Added to korrents 22 Jun 2025 · How quotes work · Something wrong? Tell us

Do you hold this korrent?Do you also believe this?

Sign in to record that you hold this, with a confidence number of your own.

Related korrents

Our reading — they may agree, disagree or merely touch the same thing. Closest first: a shared subject counts for most, then how near the wording is.

On the map

Loading the map… or open it on its own page

Open the map on its own page →