korrents

On the map

Tap a claim on the ring to put it at the centre.

← Running a model until its performance plateaus is no longer a usable…

17 connected korrents · 15 moments on record from 14 Mar 2014 to 3 Sept 2026.

Everything filed under benchmarks benchmarks Everything filed under Anthropic Anthropic Everything filed under Google Google Everything filed under scaling laws scaling laws Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Running a model until its performance plateaus is no longer a usable evaluation rule, because a well-scaffolded model keeps improving for weeks. Running a model until itsperformance plateaus is no longer ausable evaluation rule, because awell-scaffolded model keepsimproving for weeks. Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a statedbudget, or as a curve againsttest-time compute — never as a… Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it A benchmark that ranks ClaudeCode last while it stays firstin use is measuring the wrongthing, and has been for a… Last stated 4 days ago 3 Sept 2026 DR Dax Raad — holds since 2026-09-03 — tap for who they are Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it A company's staff-engineer barshould be set against the bestcompanies in the industryrather than against its own… Last stated 5 months ago 1 Apr 2026 TP Thuan Pham — holds since 2026-04-01 — tap for who they are Same subject: A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages. — tap to centre the map on it A speed difference betweenlanguages that share the LLVMbackend measures the benchmarkauthor, not the languages. Last stated a year ago 22 Mar 2025 TH ThePrimeagen — holds since 2025-03-22 — tap for who they are Same subject: AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools. — tap to centre the map on it AI learning to generate goodconjectures will never show upas a benchmark being knockeddown; it will show up as a… Last stated 2 months ago 30 Jun 2026 GS Grant Sanderson — holds since 2026-06-30 — tap for who they are Same subject: Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced. — tap to centre the map on it Benchmarks only rise onproblems somebody has alreadyframed and scored, sosaturating them does not mean… Last stated 3 months ago 24 May 2026 DS Dan Shipper — holds since 2026-05-24 — tap for who they are Same subject: Cash transfers are the index fund of development: the benchmark every actively managed aid programme should have to beat. — tap to centre the map on it Cash transfers are the indexfund of development: thebenchmark every activelymanaged aid programme should… Last stated 12 years ago 14 Mar 2014 CB Chris Blattman — holds since 2014-03-14 — tap for who they are Same subject: Every lab knows the benchmark grid is the wrong way to present a model, and publishes it anyway because everybody else does. — tap to centre the map on it Every lab knows the benchmarkgrid is the wrong way topresent a model, and publishesit anyway because everybody… Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it A lab should spend most of itscompute on research ratherthan on building the nextmodel, because research is… Last stated 6 months ago 13 Mar 2026 DP Dylan Patel — holds since 2026-03-13 — tap for who they are Same subject: After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute. — tap to centre the map on it After pre-training,post-training and test-timescaling, the fourth scalinglaw is agentic: multiplying AI… Last stated 6 months ago 23 Mar 2026 JH Jensen Huang — holds since 2026-03-23 — tap for who they are Same subject: AI model capability progress is not going to slow down soon — tap to centre the map on it AI model capability progressis not going to slow down soon Last stated 4 days ago 3 Sept 2026 AR Armin Ronacher — holds since 2026-09-03 — tap for who they are Same subject: A crewed rocket cannot be made safe by making the booster reliable, so the only real way to improve safety is to carry an escape system. — tap to centre the map on it A crewed rocket cannot be madesafe by making the boosterreliable, so the only real wayto improve safety is to carry… Last stated 3 years ago 14 Dec 2023 JB Jeff Bezos — holds since 2023-12-14 — tap for who they are Same subject: A monopolist that can no longer grow by winning new users can only grow by making its product worse for the users it already has. — tap to centre the map on it A monopolist that can nolonger grow by winning newusers can only grow by makingits product worse for the… Last stated 3 years ago 28 Jul 2023 CD Cory Doctorow — holds since 2023-07-28 — tap for who they are Same subject: A startup does not have to set out to build the greatest business in history; a merely good business is a legitimate thing to aim at. — tap to centre the map on it A startup does not have to setout to build the greatestbusiness in history; a merelygood business is a legitimate… Last stated 2 years ago 19 Jun 2024 AS Aravind Srinivas — holds since 2024-06-19 — tap for who they are Same subject: A language model asked to summarize its own system prompt risks that prompt's content biasing the summary it produces. — tap to centre the map on it A language model asked tosummarize its own systemprompt risks that prompt'scontent biasing the summary it… Last stated 5 days ago 2 Sept 2026 SW Simon Willison — holds since 2026-09-02 — tap for who they are Same subject: A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares. — tap to centre the map on it A trust that owns the missionprotects a company better thanfounder control does, which iswhy Anthropic needs no… Last stated 4 months ago 10 May 2026 ER Eric Ries — holds since 2026-05-10 — tap for who they are Same subject: Agents can build about half a million lines before the codebase dissolves into a mess, and the next model will push that to a few million. — tap to centre the map on it Agents can build about half amillion lines before thecodebase dissolves into amess, and the next model will… Last stated 6 months ago 11 Mar 2026 SY Steve Yegge — holds since 2026-03-11 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: how recently it was last stated — full and dark this week, a faint sliver at five yearsa face: someone on record holding the claim — tap it for who they are

At the centre Running a model until its performance plateaus is no longer a usable evaluation rule, because a well-scaffolded model keeps improving for weeks. Last stated 26 Jun 2026 · 2 months ago Holds Noam Brown Read this korrent →