Tap a claim on the ring to put it at the centre.
← Evals are not the durable asset people think they are: one survives…
17 connected korrents · 16 moments on record from 18 Aug 2025 to 3 Sept 2026.
Everything filed under benchmarks
benchmarks
Everything filed under reinforcement learning
reinforcement learning
Everything filed under scaling laws
scaling laws
Everything filed under TypeScript
TypeScript
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Evals are not the durable asset people think they are: one survives about one to three model generations before it saturates and has to be thrown away.
Evals are not the durable asset people think they are: one survives about one to three model generations before it saturates and has to be thrown away.
Last stated a month ago
27 Jul 2026
BC
Boris Cherny — holds since 2026-07-27 — tap for who they are
Same subject: Models look far better on evals than they are in the world because researchers, inadvertently, take inspiration from the evals when they build RL environments. — tap to centre the map on it
Models look far better on evals than they are in the world because researchers, inadvertently, take inspiration from the evals when they build RL environments.
Last stated 9 months ago
25 Nov 2025
IS
Ilya Sutskever — holds since 2025-11-25 — tap for who they are
Same subject: Before building on a gap in the general models, work out whether that gap survives six months or three years. — tap to centre the map on it
Before building on a gap in the general models, work out whether that gap survives six months or three years.
Last stated a month ago
30 Jul 2026
JD
Jeff Dean — holds since 2026-07-30 — tap for who they are
Same subject: A year from now every coding model on the market will still take a bad instruction at face value and go fix the listed issues, where a human senior engineer would insist on a rewrite. — tap to centre the map on it
A year from now every coding model on the market will still take a bad instruction at face value and go fix the listed issues, where a human senior engineer would insist on a rewrite.
Last stated 3 months ago
24 May 2026
DS
Dan Shipper — holds since 2026-05-24 — tap for who they are
Same subject: You can now generate ten thousand variants of a function faster than you could write it once, and that economics pushes software towards replacement rather than repair. — tap to centre the map on it
You can now generate ten thousand variants of a function faster than you could write it once, and that economics pushes software towards replacement rather than repair.
Last stated 4 weeks ago
12 Aug 2026
CM
Charity Majors — holds since 2026-08-12 — tap for who they are
Same subject: Running a model until its performance plateaus is no longer a usable evaluation rule, because a well-scaffolded model keeps improving for weeks. — tap to centre the map on it
Running a model until its performance plateaus is no longer a usable evaluation rule, because a well-scaffolded model keeps improving for weeks.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A programming language is a ten-year play: version one has issues, version two fixes them, version three is finally good — and only then does adoption start. — tap to centre the map on it
A programming language is a ten-year play: version one has issues, version two fixes them, version three is finally good — and only then does adoption start.
Last stated 4 months ago
13 May 2026
AH
Anders Hejlsberg — holds since 2026-05-13 — tap for who they are
Same subject: An applied AI company should not train its own foundation model, because a foundation model is the fastest deteriorating asset there is. — tap to centre the map on it
An applied AI company should not train its own foundation model, because a foundation model is the fastest deteriorating asset there is.
Last stated a year ago
18 Aug 2025
BT
Bret Taylor — holds since 2025-08-18 — tap for who they are
Same subject: Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive. — tap to centre the map on it
Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive.
Last stated 2 months ago
15 Jul 2026
DH
Dex Horthy — holds since 2026-07-15 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 4 days ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 5 months ago
1 Apr 2026
TP
Thuan Pham — holds since 2026-04-01 — tap for who they are
Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 10 months ago
17 Nov 2025
AK
Andrej Karpathy — holds since 2025-11-17 — tap for who they are
Same subject: Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago. — tap to centre the map on it
Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago.
Last stated 6 months ago
20 Mar 2026
AK
Andrej Karpathy — holds since 2026-03-20 — tap for who they are
Same subject: Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving. — tap to centre the map on it
Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving.
Last stated 11 months ago
17 Oct 2025
AK
Andrej Karpathy — holds since 2025-10-17 — tap for who they are
Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 6 months ago
13 Mar 2026
DP
Dylan Patel — holds since 2026-03-13 — tap for who they are
Same subject: After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute. — tap to centre the map on it
After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute.
Last stated 6 months ago
23 Mar 2026
JH
Jensen Huang — holds since 2026-03-23 — tap for who they are
Same subject: AI model capability progress is not going to slow down soon — tap to centre the map on it
AI model capability progress is not going to slow down soon
Last stated 4 days ago
3 Sept 2026
AR
Armin Ronacher — holds since 2026-09-03 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
Evals are not the durable asset people think they are: one survives about one to three model generations before it saturates and has to be thrown away.
Last stated 27 Jul 2026 · a month ago
Holds Boris Cherny
Read this korrent →
Similar wording
Models look far better on evals than they are in the world because researchers, inadvertently, take inspiration from the evals when they build RL environments.
Last stated 25 Nov 2025 · 9 months ago
Holds IS Ilya Sutskever
Similar wording
Before building on a gap in the general models, work out whether that gap survives six months or three years.
Last stated 30 Jul 2026 · a month ago
Holds JD Jeff Dean
Similar wording
A year from now every coding model on the market will still take a bad instruction at face value and go fix the listed issues, where a human senior engineer would insist on a rewrite.
Last stated 24 May 2026 · 3 months ago
Holds DS Dan Shipper
Similar wording
You can now generate ten thousand variants of a function faster than you could write it once, and that economics pushes software towards replacement rather than repair.
Last stated 12 Aug 2026 · 4 weeks ago
Holds CM Charity Majors
Similar wording
Running a model until its performance plateaus is no longer a usable evaluation rule, because a well-scaffolded model keeps improving for weeks.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Similar wording
A programming language is a ten-year play: version one has issues, version two fixes them, version three is finally good — and only then does adoption start.
Last stated 13 May 2026 · 4 months ago
Holds AH Anders Hejlsberg
Similar wording
An applied AI company should not train its own foundation model, because a foundation model is the fastest deteriorating asset there is.
Last stated 18 Aug 2025 · a year ago
Holds BT Bret Taylor
Similar wording
Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive.
Last stated 15 Jul 2026 · 2 months ago
Holds DH Dex Horthy
Same subject: benchmarks
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Same subject: benchmarks
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 Sept 2026 · 4 days ago
Holds Dax Raad
Same subject: benchmarks
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 1 Apr 2026 · 5 months ago
Holds TP Thuan Pham
Same subject: reinforcement learning
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 17 Nov 2025 · 10 months ago
Holds Andrej Karpathy
Same subject: reinforcement learning
Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago.
Last stated 20 Mar 2026 · 6 months ago
Holds Andrej Karpathy
Same subject: reinforcement learning
Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving.
Last stated 17 Oct 2025 · 11 months ago
Holds Andrej Karpathy
Same subject: scaling laws
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 13 Mar 2026 · 6 months ago
Holds DP Dylan Patel
Same subject: scaling laws
After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute.
Last stated 23 Mar 2026 · 6 months ago
Holds JH Jensen Huang
Same subject: scaling laws
AI model capability progress is not going to slow down soon
Last stated 3 Sept 2026 · 4 days ago
Holds Armin Ronacher