korrents

On the map

Tap a claim on the ring to put it at the centre.

← A benchmark result should be reported under a stated budget, or as a…

17 connected korrents · 15 moments on record from 14 Mar 2014 to 3 Sept 2026. Nearly all of them are about benchmarks.

Everything filed under scaling laws scaling laws Everything filed under Google Google Everything filed under Anthropic Anthropic Everything filed under compilers compilers Everything filed under mathematics mathematics Everything filed under stock market stock market Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. A benchmark result should be reportedunder a stated budget, or as a curveagainst test-time compute — never as asingle number. Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used. — tap to centre the map on it Model comparisons understate realprogress, because benchmark tablesdo not control for how muchtest-time compute each answer used. Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it A benchmark that ranks Claude Codelast while it stays first in use ismeasuring the wrong thing, and hasbeen for a year. Last stated 4 days ago 3 Sept 2026 DR Dax Raad — holds since 2026-09-03 — tap for who they are Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it A company's staff-engineer barshould be set against the bestcompanies in the industry ratherthan against its own history, whichis what makes title inflation a realcost. Last stated 5 months ago 1 Apr 2026 TP Thuan Pham — holds since 2026-04-01 — tap for who they are Same subject: A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages. — tap to centre the map on it A speed difference between languagesthat share the LLVM backend measuresthe benchmark author, not thelanguages. Last stated a year ago 22 Mar 2025 TH ThePrimeagen — holds since 2025-03-22 — tap for who they are Same subject: AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools. — tap to centre the map on it AI learning to generate goodconjectures will never show up as abenchmark being knocked down; itwill show up as a shift in howmathematicians talk about the tools. Last stated 2 months ago 30 Jun 2026 GS Grant Sanderson — holds since 2026-06-30 — tap for who they are Same subject: Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced. — tap to centre the map on it Benchmarks only rise on problemssomebody has already framed andscored, so saturating them does notmean senior engineers have beenreplaced. Last stated 3 months ago 24 May 2026 DS Dan Shipper — holds since 2026-05-24 — tap for who they are Same subject: Cash transfers are the index fund of development: the benchmark every actively managed aid programme should have to beat. — tap to centre the map on it Cash transfers are the index fund ofdevelopment: the benchmark everyactively managed aid programmeshould have to beat. Last stated 12 years ago 14 Mar 2014 CB Chris Blattman — holds since 2014-03-14 — tap for who they are Same subject: Every lab knows the benchmark grid is the wrong way to present a model, and publishes it anyway because everybody else does. — tap to centre the map on it Every lab knows the benchmark gridis the wrong way to present a model,and publishes it anyway becauseeverybody else does. Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it A lab should spend most of itscompute on research rather than onbuilding the next model, becauseresearch is where the tenfold yearlyefficiency gains come from. Last stated 6 months ago 13 Mar 2026 DP Dylan Patel — holds since 2026-03-13 — tap for who they are Same subject: After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute. — tap to centre the map on it After pre-training, post-trainingand test-time scaling, the fourthscaling law is agentic: multiplyingAI by spawning agents, and the wholeloop scales on one thing, compute. Last stated 6 months ago 23 Mar 2026 JH Jensen Huang — holds since 2026-03-23 — tap for who they are Same subject: AI model capability progress is not going to slow down soon — tap to centre the map on it AI model capability progress is notgoing to slow down soon Last stated 4 days ago 3 Sept 2026 AR Armin Ronacher — holds since 2026-09-03 — tap for who they are Same subject: Models resemble each other because pre-training is the same everywhere; what differentiates labs now is RL and post-training. — tap to centre the map on it Models resemble each other becausepre-training is the same everywhere;what differentiates labs now is RLand post-training. Last stated 9 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: One pre-trained robot model now matches or beats the specialists that were fine-tuned with reinforcement learning for the very tasks they were built for. — tap to centre the map on it One pre-trained robot model nowmatches or beats the specialiststhat were fine-tuned withreinforcement learning for the verytasks they were built for. Last stated 4 weeks ago 12 Aug 2026 CF Chelsea Finn — holds since 2026-08-12 — tap for who they are Same subject: Pre-training is a crappy evolution: the practically buildable substitute for the process that gave animals their built-in hardware. — tap to centre the map on it Pre-training is a crappy evolution:the practically buildable substitutefor the process that gave animalstheir built-in hardware. Last stated 11 months ago 17 Oct 2025 AK Andrej Karpathy — holds since 2025-10-17 — tap for who they are Same subject: Scaling laws are still working, but the models people actually use run a few months behind the maximum capability Google can deliver, because the biggest model is too slow and expensive to serve. — tap to centre the map on it Scaling laws are still working, butthe models people actually use run afew months behind the maximumcapability Google can deliver,because the biggest model is tooslow and expensive to serve. Last stated a year ago 5 Jun 2025 SP Sundar Pichai — holds since 2025-06-05 — tap for who they are Same subject: Solving one Olympiad problem with three days of Google's server time shows what is possible, but the approach does not scale. — tap to centre the map on it Solving one Olympiad problem withthree days of Google's server timeshows what is possible, but theapproach does not scale. Last stated a year ago 14 Jun 2025 TT Terence Tao — holds since 2025-06-14 — tap for who they are Same subject: A crewed rocket cannot be made safe by making the booster reliable, so the only real way to improve safety is to carry an escape system. — tap to centre the map on it A crewed rocket cannot be made safeby making the booster reliable, sothe only real way to improve safetyis to carry an escape system. Last stated 3 years ago 14 Dec 2023 JB Jeff Bezos — holds since 2023-12-14 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2014 to today (stretched back to the oldest claim here) — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. Last stated 26 Jun 2026 · 2 months ago Holds Noam Brown Read this korrent →