Tap a claim on the ring to put it at the centre.
← A published training-run cost always understates the truth, because…
8 connected korrents · 6 moments on record from 3 Feb 2025 to 30 Jul 2026.
−
+
Fit
Full text
Drag to move · pinch or scroll to zoom
Everything filed under benchmarks
benchmarks
Everything filed under scaling laws
scaling laws
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: A published training-run cost always understates the truth, because any notable model takes two to four times that run's compute in experiments alone.
A published training-run cost always understates the truth, because any notable model takes two to four times that run's compute in experiments alone.
Last stated 2 years ago
3 Feb 2025
NL
Nathan Lambert — holds since 2025-02-03 — tap for who they are
Same subject: Models will not learn to write maintainable code from today’s benchmarks, because the cost of bad architecture cannot be measured by running the unit tests — it arrives three to six months later. — tap to centre the map on it
Models will not learn to write maintainable code from today’s benchmarks, because the cost of bad architecture cannot be…
Last stated 2 months ago
15 Jul 2026
DH
Dex Horthy — holds since 2026-07-15 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a…
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: Running a model five times and keeping the best answer buys a higher benchmark score without buying a better model. — tap to centre the map on it
Running a model five times and keeping the best answer buys a higher benchmark score without buying a better model.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: Running a model until its performance plateaus is no longer a usable evaluation rule, because a well-scaffolded model keeps improving for weeks. — tap to centre the map on it
Running a model until its performance plateaus is no longer a usable evaluation rule, because a…
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used. — tap to centre the map on it
Model comparisons understate real progress, because benchmark tables do not control for how much test-time…
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it
A lab should spend most of its compute on research rather than on building the next model, because research is…
Last stated 6 months ago
13 Mar 2026
DP
Dylan Patel — holds since 2026-03-13 — tap for who they are
Same subject: Training is no longer limited by data but by compute, because most of the data models learn from is now synthetic. — tap to centre the map on it
Training is no longer limited by data but by compute, because most of the data models learn from is now…
Last stated 6 months ago
23 Mar 2026
JH
Jensen Huang — holds since 2026-03-23 — tap for who they are
Same subject: Learned approximations of slow simulators change what science is possible, by turning a six-month screening run into something you do over lunch. — tap to centre the map on it
Learned approximations of slow simulators change what science is possible, by turning a six-month screening run into…
Last stated a month ago
30 Jul 2026
JD
Jeff Dean — holds since 2026-07-30 — tap for who they are
same subject or similar wording shaded: claims sharing a subject bar: how recently it was last stated — full and dark this week, a faint sliver at five years