korrents

On the map

Tap a claim on the ring to put it at the centre.

← Evals are not the durable asset people think they are: one survives…

17 connected korrents · 16 moments on record from 18 Aug 2025 to 3 Sept 2026.

Everything filed under benchmarks benchmarks Everything filed under reinforcement learning reinforcement learning Everything filed under scaling laws scaling laws Everything filed under TypeScript TypeScript Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Evals are not the durable asset people think they are: one survives about one to three model generations before it saturates and has to be thrown away. Evals are not the durable asset peoplethink they are: one survives about one tothree model generations before itsaturates and has to be thrown away. Last stated a month ago 27 Jul 2026 BC Boris Cherny — holds since 2026-07-27 — tap for who they are Same subject: Models look far better on evals than they are in the world because researchers, inadvertently, take inspiration from the evals when they build RL environments. — tap to centre the map on it Models look far better on evals thanthey are in the world becauseresearchers, inadvertently, takeinspiration from the evals when theybuild RL environments. Last stated 9 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: Before building on a gap in the general models, work out whether that gap survives six months or three years. — tap to centre the map on it Before building on a gap in thegeneral models, work out whetherthat gap survives six months orthree years. Last stated a month ago 30 Jul 2026 JD Jeff Dean — holds since 2026-07-30 — tap for who they are Same subject: A year from now every coding model on the market will still take a bad instruction at face value and go fix the listed issues, where a human senior engineer would insist on a rewrite. — tap to centre the map on it A year from now every coding modelon the market will still take a badinstruction at face value and go fixthe listed issues, where a humansenior engineer would insist on arewrite. Last stated 3 months ago 24 May 2026 DS Dan Shipper — holds since 2026-05-24 — tap for who they are Same subject: You can now generate ten thousand variants of a function faster than you could write it once, and that economics pushes software towards replacement rather than repair. — tap to centre the map on it You can now generate ten thousandvariants of a function faster thanyou could write it once, and thateconomics pushes software towardsreplacement rather than repair. Last stated 4 weeks ago 12 Aug 2026 CM Charity Majors — holds since 2026-08-12 — tap for who they are Same subject: Running a model until its performance plateaus is no longer a usable evaluation rule, because a well-scaffolded model keeps improving for weeks. — tap to centre the map on it Running a model until itsperformance plateaus is no longer ausable evaluation rule, because awell-scaffolded model keepsimproving for weeks. Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A programming language is a ten-year play: version one has issues, version two fixes them, version three is finally good — and only then does adoption start. — tap to centre the map on it A programming language is a ten-yearplay: version one has issues,version two fixes them, versionthree is finally good — and onlythen does adoption start. Last stated 4 months ago 13 May 2026 AH Anders Hejlsberg — holds since 2026-05-13 — tap for who they are Same subject: An applied AI company should not train its own foundation model, because a foundation model is the fastest deteriorating asset there is. — tap to centre the map on it An applied AI company should nottrain its own foundation model,because a foundation model is thefastest deteriorating asset thereis. Last stated a year ago 18 Aug 2025 BT Bret Taylor — holds since 2025-08-18 — tap for who they are Same subject: Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive. — tap to centre the map on it Context engineering has stayedrelevant for a year because it isgrounded in how transformerattention works, and it will matterto anyone building on AI untilpost-transformer models arrive. Last stated 2 months ago 15 Jul 2026 DH Dex Horthy — holds since 2026-07-15 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a stated budget, oras a curve against test-time compute— never as a single number. Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it A benchmark that ranks Claude Codelast while it stays first in use ismeasuring the wrong thing, and hasbeen for a year. Last stated 4 days ago 3 Sept 2026 DR Dax Raad — holds since 2026-09-03 — tap for who they are Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it A company's staff-engineer barshould be set against the bestcompanies in the industry ratherthan against its own history, whichis what makes title inflation a realcost. Last stated 5 months ago 1 Apr 2026 TP Thuan Pham — holds since 2026-04-01 — tap for who they are Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it A verifiable task can be optimisedby reinforcement learning until aneural network performs it extremelywell. Last stated 10 months ago 17 Nov 2025 AK Andrej Karpathy — holds since 2025-11-17 — tap for who they are Same subject: Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago. — tap to centre the map on it Capability does not generalise forfree: a model that will movemountains on an agentic task stilltells the same bad joke it told fiveyears ago. Last stated 6 months ago 20 Mar 2026 AK Andrej Karpathy — holds since 2026-03-20 — tap for who they are Same subject: Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving. — tap to centre the map on it Humans barely use reinforcementlearning for intelligence — what RLthey do use goes into motor tasks,not problem solving. Last stated 11 months ago 17 Oct 2025 AK Andrej Karpathy — holds since 2025-10-17 — tap for who they are Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it A lab should spend most of itscompute on research rather than onbuilding the next model, becauseresearch is where the tenfold yearlyefficiency gains come from. Last stated 6 months ago 13 Mar 2026 DP Dylan Patel — holds since 2026-03-13 — tap for who they are Same subject: After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute. — tap to centre the map on it After pre-training, post-trainingand test-time scaling, the fourthscaling law is agentic: multiplyingAI by spawning agents, and the wholeloop scales on one thing, compute. Last stated 6 months ago 23 Mar 2026 JH Jensen Huang — holds since 2026-03-23 — tap for who they are Same subject: AI model capability progress is not going to slow down soon — tap to centre the map on it AI model capability progress is notgoing to slow down soon Last stated 4 days ago 3 Sept 2026 AR Armin Ronacher — holds since 2026-09-03 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre Evals are not the durable asset people think they are: one survives about one to three model generations before it saturates and has to be thrown away. Last stated 27 Jul 2026 · a month ago Holds Boris Cherny Read this korrent →