What Sebastian Ruder thinks about benchmarks
Research scientist at Meta working on multilingual models and evaluation; led the multilingual team at Cohere and was at Google DeepMind before that.
Everything they publish, on ppll ↗
Sebastian Ruder did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
2 dated positions, 2024, in their own words. Our reading of what Sebastian Ruder has said — not written or endorsed by them.
-
Their wordsGPT models perform much better on coding problems released before their pre-training data cut-off.
↗The Evolving Landscape of LLM Evaluationruder.io 1st of 2 in this piece
-
Their wordsThe time when benchmarks lasted multiple decades has passed. Going forward, we will rely less on public benchmark results.
↗The Evolving Landscape of LLM Evaluationruder.io 2nd of 2 in this piece