korrents

On the map

Tap a claim on the ring to put it at the centre.

← A test meant to find the best candidates must make perfect scores…

12 connected korrents · 10 moments on record from 17 Apr 2012 to 3 Sept 2026.

Everything filed under benchmarks benchmarks Everything filed under software quality software quality Everything filed under writing writing Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: A test meant to find the best candidates must make perfect scores vanishingly rare, so that the top of the distribution can still be told apart. A test meant to find the best candidatesmust make perfect scores vanishingly rare,so that the top of the distribution canstill be told apart. Last stated a year ago 15 Apr 2025 BC Bryan Caplan — holds since 2025-04-15 — tap for who they are Same subject: Elite programmes stay with tests that cannot separate the best candidates because doing so is a stealthy move in a wider war against merit. — tap to centre the map on it Elite programmes stay with teststhat cannot separate the bestcandidates because doing so is astealthy move in a wider war againstmerit. Last stated a year ago 15 Apr 2025 BC Bryan Caplan — holds since 2025-04-15 — tap for who they are Same subject: Running a model five times and keeping the best answer buys a higher benchmark score without buying a better model. — tap to centre the map on it Running a model five times andkeeping the best answer buys ahigher benchmark score withoutbuying a better model. Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: Making test coverage a target destroys it, because high coverage numbers are easy to reach with worthless tests. — tap to centre the map on it Making test coverage a targetdestroys it, because high coveragenumbers are easy to reach withworthless tests. Last stated 14 years ago 17 Apr 2012 MF Martin Fowler — holds since 2012-04-17 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a stated budget, oras a curve against test-time compute— never as a single number. Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it A benchmark that ranks Claude Codelast while it stays first in use ismeasuring the wrong thing, and hasbeen for a year. Last stated 5 days ago 3 Sept 2026 DR Dax Raad — holds since 2026-09-03 — tap for who they are Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it A company's staff-engineer barshould be set against the bestcompanies in the industry ratherthan against its own history, whichis what makes title inflation a realcost. Last stated 5 months ago 1 Apr 2026 TP Thuan Pham — holds since 2026-04-01 — tap for who they are Same subject: A language-agnostic conformance suite is the most powerful thing you can hand a coding agent, because the whole instruction becomes: write code until these tests pass. — tap to centre the map on it A language-agnostic conformancesuite is the most powerful thing youcan hand a coding agent, because thewhole instruction becomes: writecode until these tests pass. Last stated 6 months ago 19 Mar 2026 SW Simon Willison — holds since 2026-03-19 — tap for who they are Same subject: A passing test suite is not evidence that the software works, so an agent must also be made to start the thing and exercise it the way a person would. — tap to centre the map on it A passing test suite is not evidencethat the software works, so an agentmust also be made to start the thingand exercise it the way a personwould. Last stated 6 months ago 19 Mar 2026 SW Simon Willison — holds since 2026-03-19 — tap for who they are Same subject: Adding a config option so nobody breaks compounds into permutations no test suite can cover, because it is infinitely harder to evolve software that has users. — tap to centre the map on it Adding a config option so nobodybreaks compounds into permutationsno test suite can cover, because itis infinitely harder to evolvesoftware that has users. Last stated 4 weeks ago 10 Aug 2026 PS Peter Steinberger — holds since 2026-08-10 — tap for who they are Same subject: A browser that answers with generated prose instead of links is not a web browser but an anti-web browser, and it hides that substitution from the person using it. — tap to centre the map on it A browser that answers withgenerated prose instead of links isnot a web browser but an anti-webbrowser, and it hides thatsubstitution from the person usingit. Last stated 11 months ago 22 Oct 2025 AD Anil Dash — holds since 2025-10-22 — tap for who they are Same subject: A first draft should not be graded good or bad; it is only the material that taste then gets to act on. — tap to centre the map on it A first draft should not be gradedgood or bad; it is only the materialthat taste then gets to act on. Last stated a month ago 6 Aug 2026 GS George Saunders — holds since 2026-08-06 — tap for who they are Same subject: A plain text file and a basic editor are everything anyone needs in order to think and write well; the extra tools add nothing. — tap to centre the map on it A plain text file and a basic editorare everything anyone needs in orderto think and write well; the extratools add nothing. Last stated 5 years ago 2 Mar 2022 DS Derek Sivers — holds since 2022-03-02 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2012 to today (stretched back to the oldest claim here) — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre A test meant to find the best candidates must make perfect scores vanishingly rare, so that the top of the distribution can still be told apart. Last stated 15 Apr 2025 · a year ago Holds Bryan Caplan Read this korrent →