Tap a claim on the ring to put it at the centre.
← A benchmark result should be reported under a stated budget, or as a…
17 connected korrents · 15 moments on record from 14 Mar 2014 to 3 Sept 2026. Nearly all of them are about benchmarks .
Everything filed under scaling laws
scaling laws
Everything filed under Google
Google
Everything filed under Anthropic
Anthropic
Everything filed under compilers
compilers
Everything filed under mathematics
mathematics
Everything filed under stock market
stock market
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used. — tap to centre the map on it
Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 4 days ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 5 months ago
1 Apr 2026
TP
Thuan Pham — holds since 2026-04-01 — tap for who they are
Same subject: A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages. — tap to centre the map on it
A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages.
Last stated a year ago
22 Mar 2025
TH
ThePrimeagen — holds since 2025-03-22 — tap for who they are
Same subject: AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools. — tap to centre the map on it
AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools.
Last stated 2 months ago
30 Jun 2026
GS
Grant Sanderson — holds since 2026-06-30 — tap for who they are
Same subject: Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced. — tap to centre the map on it
Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced.
Last stated 3 months ago
24 May 2026
DS
Dan Shipper — holds since 2026-05-24 — tap for who they are
Same subject: Cash transfers are the index fund of development: the benchmark every actively managed aid programme should have to beat. — tap to centre the map on it
Cash transfers are the index fund of development: the benchmark every actively managed aid programme should have to beat.
Last stated 12 years ago
14 Mar 2014
CB
Chris Blattman — holds since 2014-03-14 — tap for who they are
Same subject: Every lab knows the benchmark grid is the wrong way to present a model, and publishes it anyway because everybody else does. — tap to centre the map on it
Every lab knows the benchmark grid is the wrong way to present a model, and publishes it anyway because everybody else does.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 6 months ago
13 Mar 2026
DP
Dylan Patel — holds since 2026-03-13 — tap for who they are
Same subject: After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute. — tap to centre the map on it
After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute.
Last stated 6 months ago
23 Mar 2026
JH
Jensen Huang — holds since 2026-03-23 — tap for who they are
Same subject: AI model capability progress is not going to slow down soon — tap to centre the map on it
AI model capability progress is not going to slow down soon
Last stated 4 days ago
3 Sept 2026
AR
Armin Ronacher — holds since 2026-09-03 — tap for who they are
Same subject: Models resemble each other because pre-training is the same everywhere; what differentiates labs now is RL and post-training. — tap to centre the map on it
Models resemble each other because pre-training is the same everywhere; what differentiates labs now is RL and post-training.
Last stated 9 months ago
25 Nov 2025
IS
Ilya Sutskever — holds since 2025-11-25 — tap for who they are
Same subject: One pre-trained robot model now matches or beats the specialists that were fine-tuned with reinforcement learning for the very tasks they were built for. — tap to centre the map on it
One pre-trained robot model now matches or beats the specialists that were fine-tuned with reinforcement learning for the very tasks they were built for.
Last stated 4 weeks ago
12 Aug 2026
CF
Chelsea Finn — holds since 2026-08-12 — tap for who they are
Same subject: Pre-training is a crappy evolution: the practically buildable substitute for the process that gave animals their built-in hardware. — tap to centre the map on it
Pre-training is a crappy evolution: the practically buildable substitute for the process that gave animals their built-in hardware.
Last stated 11 months ago
17 Oct 2025
AK
Andrej Karpathy — holds since 2025-10-17 — tap for who they are
Same subject: Scaling laws are still working, but the models people actually use run a few months behind the maximum capability Google can deliver, because the biggest model is too slow and expensive to serve. — tap to centre the map on it
Scaling laws are still working, but the models people actually use run a few months behind the maximum capability Google can deliver, because the biggest model is too slow and expensive to serve.
Last stated a year ago
5 Jun 2025
SP
Sundar Pichai — holds since 2025-06-05 — tap for who they are
Same subject: Solving one Olympiad problem with three days of Google's server time shows what is possible, but the approach does not scale. — tap to centre the map on it
Solving one Olympiad problem with three days of Google's server time shows what is possible, but the approach does not scale.
Last stated a year ago
14 Jun 2025
TT
Terence Tao — holds since 2025-06-14 — tap for who they are
Same subject: A crewed rocket cannot be made safe by making the booster reliable, so the only real way to improve safety is to carry an escape system. — tap to centre the map on it
A crewed rocket cannot be made safe by making the booster reliable, so the only real way to improve safety is to carry an escape system.
Last stated 3 years ago
14 Dec 2023
JB
Jeff Bezos — holds since 2023-12-14 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2014 to today (stretched back to the oldest claim here) — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Read this korrent →
Same subject: benchmarks, scaling laws
Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Same subject: benchmarks
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 Sept 2026 · 4 days ago
Holds Dax Raad
Same subject: benchmarks
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 1 Apr 2026 · 5 months ago
Holds TP Thuan Pham
Same subject: benchmarks
A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages.
Last stated 22 Mar 2025 · a year ago
Holds TH ThePrimeagen
Same subject: benchmarks
AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools.
Last stated 30 Jun 2026 · 2 months ago
Holds GS Grant Sanderson
Same subject: benchmarks
Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced.
Last stated 24 May 2026 · 3 months ago
Holds DS Dan Shipper
Same subject: benchmarks
Cash transfers are the index fund of development: the benchmark every actively managed aid programme should have to beat.
Last stated 14 Mar 2014 · 12 years ago
Holds CB Chris Blattman
Same subject: benchmarks
Every lab knows the benchmark grid is the wrong way to present a model, and publishes it anyway because everybody else does.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Same subject: scaling laws
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 13 Mar 2026 · 6 months ago
Holds DP Dylan Patel
Same subject: scaling laws
After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute.
Last stated 23 Mar 2026 · 6 months ago
Holds JH Jensen Huang
Same subject: scaling laws
AI model capability progress is not going to slow down soon
Last stated 3 Sept 2026 · 4 days ago
Holds Armin Ronacher
Same subject: reinforcement learning
Models resemble each other because pre-training is the same everywhere; what differentiates labs now is RL and post-training.
Last stated 25 Nov 2025 · 9 months ago
Holds IS Ilya Sutskever
Same subject: reinforcement learning
One pre-trained robot model now matches or beats the specialists that were fine-tuned with reinforcement learning for the very tasks they were built for.
Last stated 12 Aug 2026 · 4 weeks ago
Holds CF Chelsea Finn
Same subject: reinforcement learning
Pre-training is a crappy evolution: the practically buildable substitute for the process that gave animals their built-in hardware.
Last stated 17 Oct 2025 · 11 months ago
Holds Andrej Karpathy
Same subject: Google
Scaling laws are still working, but the models people actually use run a few months behind the maximum capability Google can deliver, because the biggest model is too slow and expensive to serve.
Last stated 5 Jun 2025 · a year ago
Holds Sundar Pichai
Same subject: Google
Solving one Olympiad problem with three days of Google's server time shows what is possible, but the approach does not scale.
Last stated 14 Jun 2025 · a year ago
Holds TT Terence Tao
Same subject: Google
A crewed rocket cannot be made safe by making the booster reliable, so the only real way to improve safety is to carry an escape system.
Last stated 14 Dec 2023 · 3 years ago
Holds JB Jeff Bezos