Tap a claim on the ring to put it at the centre.
← Evaluating a model properly would mean delaying its release, and competitive pressure means no lab will.
17 connected korrents · 14 moments on record from 17 Oct 2025 to 4 Sept 2026.
Everything filed under benchmarks
benchmarks
Everything filed under reinforcement learning
reinforcement learning
Everything filed under scaling laws
scaling laws
Everything filed under Anthropic
Anthropic
Everything filed under design
design
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Evaluating a model properly would mean delaying its release, and competitive pressure means no lab will.
Evaluating a model properly would mean delaying its release, and competitive pressure means no lab will.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: Models look far better on evals than they are in the world because researchers, inadvertently, take inspiration from the evals when they build RL environments. — tap to centre the map on it
Models look far better on evals than they are in the world because researchers, inadvertently, take inspiration from the evals when they build RL environments.
Last stated 9 months ago
25 Nov 2025
IS
Ilya Sutskever — holds since 2025-11-25 — tap for who they are
Same subject: When a capability evaluation benchmark is saturated without real-world verification, a model should be presumed to have that capability until shown otherwise. — tap to centre the map on it
When a capability evaluation benchmark is saturated without real-world verification, a model should be presumed to have that capability until shown otherwise.
Last stated 4 days ago
4 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-04 — tap for who they are
Same subject: Running a model until its performance plateaus is no longer a usable evaluation rule, because a well-scaffolded model keeps improving for weeks. — tap to centre the map on it
Running a model until its performance plateaus is no longer a usable evaluation rule, because a well-scaffolded model keeps improving for weeks.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: Today's models are the student who drilled 10,000 hours for the programming contest, which is exactly why what they learn does not carry anywhere else. — tap to centre the map on it
Today's models are the student who drilled 10,000 hours for the programming contest, which is exactly why what they learn does not carry anywhere else.
Last stated 9 months ago
25 Nov 2025
IS
Ilya Sutskever — holds since 2025-11-25 — tap for who they are
Same subject: The model is table stakes; evaluation and metrics are the strategic moat, and they should be built before the technology. — tap to centre the map on it
The model is table stakes; evaluation and metrics are the strategic moat, and they should be built before the technology.
Last stated a month ago
3 Aug 2026
DD
Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
Same subject: A narrow, highly accurate model can beat a general one inside its own domain, as AlphaFold did, and materials science and chip design are next. — tap to centre the map on it
A narrow, highly accurate model can beat a general one inside its own domain, as AlphaFold did, and materials science and chip design are next.
Last stated a month ago
30 Jul 2026
JD
Jeff Dean — holds since 2026-07-30 — tap for who they are
Same subject: Running a model five times and keeping the best answer buys a higher benchmark score without buying a better model. — tap to centre the map on it
Running a model five times and keeping the best answer buys a higher benchmark score without buying a better model.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 10 months ago
17 Nov 2025
AK
Andrej Karpathy — holds since 2025-11-17 — tap for who they are
Same subject: Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago. — tap to centre the map on it
Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago.
Last stated 6 months ago
20 Mar 2026
AK
Andrej Karpathy — holds since 2026-03-20 — tap for who they are
Same subject: Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving. — tap to centre the map on it
Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving.
Last stated 11 months ago
17 Oct 2025
AK
Andrej Karpathy — holds since 2025-10-17 — tap for who they are
Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 6 months ago
13 Mar 2026
DP
Dylan Patel — holds since 2026-03-13 — tap for who they are
Same subject: After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute. — tap to centre the map on it
After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute.
Last stated 6 months ago
23 Mar 2026
JH
Jensen Huang — holds since 2026-03-23 — tap for who they are
Same subject: AI model capability progress is not going to slow down soon — tap to centre the map on it
AI model capability progress is not going to slow down soon
Last stated 5 days ago
3 Sept 2026
AR
Armin Ronacher — holds since 2026-09-03 — tap for who they are
Same subject: A language model asked to summarize its own system prompt risks that prompt's content biasing the summary it produces. — tap to centre the map on it
A language model asked to summarize its own system prompt risks that prompt's content biasing the summary it produces.
Last stated 6 days ago
2 Sept 2026
SW
Simon Willison — holds since 2026-09-02 — tap for who they are
Same subject: A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares. — tap to centre the map on it
A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares.
Last stated 4 months ago
10 May 2026
ER
Eric Ries — holds since 2026-05-10 — tap for who they are
Same subject: Agents can build about half a million lines before the codebase dissolves into a mess, and the next model will push that to a few million. — tap to centre the map on it
Agents can build about half a million lines before the codebase dissolves into a mess, and the next model will push that to a few million.
Last stated 6 months ago
11 Mar 2026
SY
Steve Yegge — holds since 2026-03-11 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
Evaluating a model properly would mean delaying its release, and competitive pressure means no lab will.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Read this korrent →
Similar wording
Models look far better on evals than they are in the world because researchers, inadvertently, take inspiration from the evals when they build RL environments.
Last stated 25 Nov 2025 · 9 months ago
Holds IS Ilya Sutskever
Similar wording
When a capability evaluation benchmark is saturated without real-world verification, a model should be presumed to have that capability until shown otherwise.
Last stated 4 Sept 2026 · 4 days ago
Holds ZM Zvi Mowshowitz
Similar wording
Running a model until its performance plateaus is no longer a usable evaluation rule, because a well-scaffolded model keeps improving for weeks.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Similar wording
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Similar wording
Today's models are the student who drilled 10,000 hours for the programming contest, which is exactly why what they learn does not carry anywhere else.
Last stated 25 Nov 2025 · 9 months ago
Holds IS Ilya Sutskever
Similar wording
The model is table stakes; evaluation and metrics are the strategic moat, and they should be built before the technology.
Last stated 3 Aug 2026 · a month ago
Holds DD Dmitri Dolgov
Similar wording
A narrow, highly accurate model can beat a general one inside its own domain, as AlphaFold did, and materials science and chip design are next.
Last stated 30 Jul 2026 · a month ago
Holds JD Jeff Dean
Similar wording
Running a model five times and keeping the best answer buys a higher benchmark score without buying a better model.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Same subject: reinforcement learning
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 17 Nov 2025 · 10 months ago
Holds Andrej Karpathy
Same subject: reinforcement learning
Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago.
Last stated 20 Mar 2026 · 6 months ago
Holds Andrej Karpathy
Same subject: reinforcement learning
Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving.
Last stated 17 Oct 2025 · 11 months ago
Holds Andrej Karpathy
Same subject: scaling laws
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 13 Mar 2026 · 6 months ago
Holds DP Dylan Patel
Same subject: scaling laws
After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute.
Last stated 23 Mar 2026 · 6 months ago
Holds JH Jensen Huang
Same subject: scaling laws
AI model capability progress is not going to slow down soon
Last stated 3 Sept 2026 · 5 days ago
Holds Armin Ronacher
Same subject: Anthropic
A language model asked to summarize its own system prompt risks that prompt's content biasing the summary it produces.
Last stated 2 Sept 2026 · 6 days ago
Holds Simon Willison
Same subject: Anthropic
A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares.
Last stated 10 May 2026 · 4 months ago
Holds ER Eric Ries
Same subject: Anthropic
Agents can build about half a million lines before the codebase dissolves into a mess, and the next model will push that to a few million.
Last stated 11 Mar 2026 · 6 months ago
Holds SY Steve Yegge