Tap a claim on the ring to put it at the centre.
← Language models perform better on benchmark problems released before…
17 connected korrents · 17 moments on record from 11 Oct 2018 to 3 Sept 2026.
Everything filed under benchmarks
benchmarks
Everything filed under scaling laws
scaling laws
Everything filed under Anthropic
Anthropic
Everything filed under measuring intelligence
measuring intelligence
Everything filed under reinforcement learning
reinforcement learning
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores.
Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores.
Last stated 2 years ago
13 May 2024
SR
Sebastian Ruder — holds since 2024-05-13 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used. — tap to centre the map on it
Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: The right bar for an LLM evaluator is human-level performance, not perfect accuracy. — tap to centre the map on it
The right bar for an LLM evaluator is human-level performance, not perfect accuracy.
Last stated 10 months ago
23 Nov 2025
EY
Eugene Yan — holds since 2025-11-23 — tap for who they are
Same subject: Current techniques restrict pre-trained representation power because standard language models are unidirectional. — tap to centre the map on it
Current techniques restrict pre-trained representation power because standard language models are unidirectional.
Last stated 8 years ago
11 Oct 2018
KT
Kristina Toutanova — holds since 2018-10-11 — tap for who they are
MC
Ming-Wei Chang — holds since 2018-10-11 — tap for who they are
KL
Kenton Lee — holds since 2018-10-11 — tap for who they are
JD
Jacob Devlin — holds since 2018-10-11 — tap for who they are
Same subject: Large language models could still plateau, and that possibility should be held open even though no evidence of it has appeared. — tap to centre the map on it
Large language models could still plateau, and that possibility should be held open even though no evidence of it has appeared.
Last stated 4 weeks ago
26 Aug 2026
DH
David Heinemeier Hansson — holds since 2026-08-26 — tap for who they are
Same subject: The LLM line of research will reach a capability plateau. — tap to centre the map on it
The LLM line of research will reach a capability plateau.
Last stated a month ago
7 Aug 2026
FC
François Chollet — no longer holds since 2026-08-07 — tap for who they are
GM
Gary Marcus — holds since 2025-06-07 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 weeks ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 6 months ago
1 Apr 2026
TP
Thuan Pham — holds since 2026-04-01 — tap for who they are
Same subject: A true artificial general intelligence cannot exist without being recognized as a moral subject. — tap to centre the map on it
A true artificial general intelligence cannot exist without being recognized as a moral subject.
Last stated a year ago
10 Jun 2025
SH
Samuel Hammond — holds since 2025-06-10 — tap for who they are
Same subject: Both the special-purpose-programs view and the blank-slate view of human intelligence are likely incorrect. — tap to centre the map on it
Both the special-purpose-programs view and the blank-slate view of human intelligence are likely incorrect.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: Every generation's test for machine intelligence is whatever machines cannot do yet, so in twenty years the question will be whether it can reproduce. — tap to centre the map on it
Every generation's test for machine intelligence is whatever machines cannot do yet, so in twenty years the question will be whether it can reproduce.
Last stated 3 years ago
29 Jun 2023
GH
George Hotz — holds since 2023-06-29 — tap for who they are
Same subject: A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel. — tap to centre the map on it
A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel.
Last stated 2 months ago
3 Aug 2026
DD
Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 10 months ago
17 Nov 2025
AK
Andrej Karpathy — holds since 2025-11-17 — tap for who they are
Same subject: Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model. — tap to centre the map on it
Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model.
Last stated 2 months ago
3 Aug 2026
DD
Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
Same subject: A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares. — tap to centre the map on it
A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares.
Last stated 4 months ago
10 May 2026
ER
Eric Ries — holds since 2026-05-10 — tap for who they are
Same subject: Agents can build about half a million lines before the codebase dissolves into a mess, and the next model will push that to a few million. — tap to centre the map on it
Agents can build about half a million lines before the codebase dissolves into a mess, and the next model will push that to a few million.
Last stated 6 months ago
11 Mar 2026
SY
Steve Yegge — holds since 2026-03-11 — tap for who they are
Same subject: AI agents and AI coding will run on servers and from the cloud first, not on your laptop. — tap to centre the map on it
AI agents and AI coding will run on servers and from the cloud first, not on your laptop.
Last stated 3 months ago
28 Jun 2026
PL
Pieter Levels — holds since 2026-06-28 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone on record holding the claim — tap it for who they are faded, dashed ring: they no longer hold it — they changed their mind
At the centre
Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores.
Last stated 13 May 2024 · 2 years ago
Holds Sebastian Ruder
Read this korrent →
Same subject: benchmarks, scaling laws
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 26 Jun 2026 · 3 months ago
Holds Noam Brown
Same subject: benchmarks, scaling laws
Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used.
Last stated 26 Jun 2026 · 3 months ago
Holds Noam Brown
Same subject: LLMs, benchmarks
The right bar for an LLM evaluator is human-level performance, not perfect accuracy.
Last stated 23 Nov 2025 · 10 months ago
Holds Eugene Yan
Same subject: LLMs, scaling laws
Current techniques restrict pre-trained representation power because standard language models are unidirectional.
Last stated 11 Oct 2018 · 8 years ago
Holds Kristina Toutanova Ming-Wei Chang Kenton Lee Jacob Devlin
Same subject: LLMs, scaling laws
Large language models could still plateau, and that possibility should be held open even though no evidence of it has appeared.
Last stated 26 Aug 2026 · 4 weeks ago
Holds David Heinemeier Hansson
Same subject: LLMs, scaling laws
The LLM line of research will reach a capability plateau.
Last stated 7 Aug 2026 · a month ago
Holds Gary MarcusNo longer holds François Chollet
Same subject: benchmarks
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 Sept 2026 · 3 weeks ago
Holds Dax Raad
Same subject: benchmarks
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 1 Apr 2026 · 6 months ago
Holds Thuan Pham
Same subject: measuring intelligence
A true artificial general intelligence cannot exist without being recognized as a moral subject.
Last stated 10 Jun 2025 · a year ago
Holds Samuel Hammond
Same subject: measuring intelligence
Both the special-purpose-programs view and the blank-slate view of human intelligence are likely incorrect.
Last stated 5 Nov 2019 · 7 years ago
Holds François Chollet
Same subject: measuring intelligence
Every generation's test for machine intelligence is whatever machines cannot do yet, so in twenty years the question will be whether it can reproduce.
Last stated 29 Jun 2023 · 3 years ago
Holds George Hotz
Same subject: reinforcement learning
A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel.
Last stated 3 Aug 2026 · 2 months ago
Holds Dmitri Dolgov
Same subject: reinforcement learning
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 17 Nov 2025 · 10 months ago
Holds Andrej Karpathy
Same subject: reinforcement learning
Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model.
Last stated 3 Aug 2026 · 2 months ago
Holds Dmitri Dolgov
Same subject: Anthropic
A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares.
Last stated 10 May 2026 · 4 months ago
Holds Eric Ries
Same subject: Anthropic
Agents can build about half a million lines before the codebase dissolves into a mess, and the next model will push that to a few million.
Last stated 11 Mar 2026 · 6 months ago
Holds Steve Yegge
Same subject: Anthropic
AI agents and AI coding will run on servers and from the cloud first, not on your laptop.
Last stated 28 Jun 2026 · 3 months ago
Holds Pieter Levels