korrents

On the map

Tap a claim on the ring to put it at the centre.

← Model comparisons understate real progress, because benchmark tables…

8 connected korrents · 7 moments on record from 14 Mar 2014 to 3 Sept 2026.

Everything filed under benchmarks benchmarks Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used. Model comparisons understate realprogress, because benchmark tablesdo not control for how muchtest-time compute each answer used. NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a statedbudget, or as a curve againsttest-time compute — never as a… NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it A benchmark that ranks ClaudeCode last while it stays firstin use is measuring the wrongthing, and has been for a… DR Dax Raad — holds since 2026-09-03 — tap for who they are Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it A company's staff-engineer barshould be set against the bestcompanies in the industryrather than against its own… TP Thuan Pham — holds since 2026-04-01 — tap for who they are Same subject: A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages. — tap to centre the map on it A speed difference betweenlanguages that share the LLVMbackend measures the benchmarkauthor, not the languages. TH ThePrimeagen — holds since 2025-03-22 — tap for who they are Same subject: AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools. — tap to centre the map on it AI learning to generate goodconjectures will never show upas a benchmark being knockeddown; it will show up as a… GS Grant Sanderson — holds since 2026-06-30 — tap for who they are Same subject: Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced. — tap to centre the map on it Benchmarks only rise onproblems somebody has alreadyframed and scored, sosaturating them does not mean… DS Dan Shipper — holds since 2026-05-24 — tap for who they are Same subject: Cash transfers are the index fund of development: the benchmark every actively managed aid programme should have to beat. — tap to centre the map on it Cash transfers are the indexfund of development: thebenchmark every activelymanaged aid programme should… CB Chris Blattman — holds since 2014-03-14 — tap for who they are Same subject: Every lab knows the benchmark grid is the wrong way to present a model, and publishes it anyway because everybody else does. — tap to centre the map on it Every lab knows the benchmarkgrid is the wrong way topresent a model, and publishesit anyway because everybody… NB Noam Brown — holds since 2026-06-26 — tap for who they are
same subject or similar wordingshaded: claims sharing a subject

At the centre Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used. Holds Noam Brown Read this korrent →