Tap a claim on the ring to put it at the centre.
← Judging one model worse than another from a one-shot, probabilistic…
17 connected korrents · 14 moments from 11 Jun 2021 to 27 Sept 2026.
Everything filed under LLMs
LLMs
Everything filed under benchmarks
benchmarks
Everything filed under AI writing
AI writing
Everything filed under open source
open source
Everything filed under OpenAI
OpenAI
Everything filed under Anthropic
Anthropic
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Judging one model worse than another from a one-shot, probabilistic output on a given day is not a good way to evaluate it.
Judging one model worse than another from a one-shot, probabilistic output on a given day is not a good way to evaluate it.
Last stated 5 months ago
19 May 2026
TG
Tommy Geoco — holds since 2026-05-19 — tap for who they are
Same subject: Trusting model-written code you have not read became defensible only with Opus 4.5, and only for classes of problem you have already watched a model handle. — tap to centre the map on it
Trusting model-written code you have not read became defensible only with Opus 4.5, and only for classes of problem you have already watched a model handle.
Last stated 7 months ago
19 Mar 2026
SW
Simon Willison — holds since 2026-03-19 — tap for who they are
Same subject: Running a model five times and keeping the best answer buys a higher benchmark score without buying a better model. — tap to centre the map on it
Running a model five times and keeping the best answer buys a higher benchmark score without buying a better model.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores. — tap to centre the map on it
Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores.
Last stated 2 years ago
13 May 2024
SR
Sebastian Ruder — holds since 2024-05-13 — tap for who they are
Same subject: Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used. — tap to centre the map on it
Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: Claude Opus 5.5 is a Tier 2 cyber model, despite Anthropic placing it in the lower category. — tap to centre the map on it
Claude Opus 5.5 is a Tier 2 cyber model, despite Anthropic placing it in the lower category.
Last stated 2 weeks ago
23 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-23 — tap for who they are
Same subject: The release of Opus 4.5 in November 2025 was the dividing line, because it was the first time an agent produced work uncannily close to what its user would have written themselves. — tap to centre the map on it
The release of Opus 4.5 in November 2025 was the dividing line, because it was the first time an agent produced work uncannily close to what its user would have written themselves.
Last stated a month ago
26 Aug 2026
DH
David Heinemeier Hansson — holds since 2026-08-26 — tap for who they are
Same subject: Evaluating a model properly would mean delaying its release, and competitive pressure means no lab will. — tap to centre the map on it
Evaluating a model properly would mean delaying its release, and competitive pressure means no lab will.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: Criticism that a simpler, scaled-up architecture underperforms a heavily hand-optimized baseline actually demonstrates why that architecture-scaling research matters. — tap to centre the map on it
Criticism that a simpler, scaled-up architecture underperforms a heavily hand-optimized baseline actually demonstrates why that architecture-scaling research matters.
Last stated 5 years ago
11 Jun 2021
GW
Gwern — holds since 2021-06-11 — tap for who they are
Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it
A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests.
Last stated 3 years ago
18 Mar 2024
SA
Sam Altman — holds since 2024-03-18 — tap for who they are
Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it
Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers.
Last stated 3 years ago
18 Mar 2024
SA
Sam Altman — holds since 2024-03-18 — tap for who they are
Same subject: Agents are going to be incorporated into mainstream interfaces — tap to centre the map on it
Agents are going to be incorporated into mainstream interfaces
Last stated 6 months ago
25 Mar 2026
DH
David Heinemeier Hansson — holds since 2026-03-25 — tap for who they are
Same subject: A CEO who sends out an AI-written strategy memo is modelling that it is fine to outsource thinking and strategy. — tap to centre the map on it
A CEO who sends out an AI-written strategy memo is modelling that it is fine to outsource thinking and strategy.
Last stated a week ago
27 Sept 2026
MG
Molly Graham — holds since 2026-09-27 — tap for who they are
Same subject: A stream of AI-written papers with any error rate at all becomes insufferable, because finding the error costs more than the paper is worth even at ninety-nine percent. — tap to centre the map on it
A stream of AI-written papers with any error rate at all becomes insufferable, because finding the error costs more than the paper is worth even at ninety-nine percent.
Last stated 3 months ago
30 Jun 2026
GS
Grant Sanderson — holds since 2026-06-30 — tap for who they are
Same subject: AI can now produce in minutes work that used to take weeks to build. — tap to centre the map on it
AI can now produce in minutes work that used to take weeks to build.
Last stated 4 months ago
26 May 2026
CB
Carlos Alexandro Becker — holds since 2026-05-26 — tap for who they are
Same subject: A business making hundreds of millions from an open source project owes it money or contribution, and can be refused its trademarks if it gives neither. — tap to centre the map on it
A business making hundreds of millions from an open source project owes it money or contribution, and can be refused its trademarks if it gives neither.
Last stated 2 years ago
26 Sept 2024
MM
Matt Mullenweg — holds since 2024-09-26 — tap for who they are
Same subject: A codec does not really exist until everyone can decode it. — tap to centre the map on it
A codec does not really exist until everyone can decode it.
Last stated 4 months ago
31 May 2026
JK
Jean-Baptiste Kempf — holds since 2026-05-31 — tap for who they are
Same subject: A company that will not accept a project’s terms is free to go and use a more permissive project instead. — tap to centre the map on it
A company that will not accept a project’s terms is free to go and use a more permissive project instead.
Last stated 2 years ago
26 Sept 2024
MM
Matt Mullenweg — holds since 2024-09-26 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone who holds the claim — tap it for who they are
At the centre
Judging one model worse than another from a one-shot, probabilistic output on a given day is not a good way to evaluate it.
Last stated 19 May 2026 · 5 months ago
Holds Tommy Geoco
Read this korrent →
Similar wording
Trusting model-written code you have not read became defensible only with Opus 4.5, and only for classes of problem you have already watched a model handle.
Last stated 19 Mar 2026 · 7 months ago
Holds Simon Willison
Similar wording
Running a model five times and keeping the best answer buys a higher benchmark score without buying a better model.
Last stated 26 Jun 2026 · 3 months ago
Holds Noam Brown
Similar wording
Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores.
Last stated 13 May 2024 · 2 years ago
Holds Sebastian Ruder
Similar wording
Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used.
Last stated 26 Jun 2026 · 3 months ago
Holds Noam Brown
Similar wording
Claude Opus 5.5 is a Tier 2 cyber model, despite Anthropic placing it in the lower category.
Last stated 23 Sept 2026 · 2 weeks ago
Holds Zvi Mowshowitz
Similar wording
The release of Opus 4.5 in November 2025 was the dividing line, because it was the first time an agent produced work uncannily close to what its user would have written themselves.
Last stated 26 Aug 2026 · a month ago
Holds David Heinemeier Hansson
Similar wording
Evaluating a model properly would mean delaying its release, and competitive pressure means no lab will.
Last stated 26 Jun 2026 · 3 months ago
Holds Noam Brown
Similar wording
Criticism that a simpler, scaled-up architecture underperforms a heavily hand-optimized baseline actually demonstrates why that architecture-scaling research matters.
Last stated 11 Jun 2021 · 5 years ago
Holds Gwern
Same subject: OpenAI
A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests.
Last stated 18 Mar 2024 · 3 years ago
Holds Sam Altman
Same subject: OpenAI
Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers.
Last stated 18 Mar 2024 · 3 years ago
Holds Sam Altman
Same subject: OpenAI
Agents are going to be incorporated into mainstream interfaces
Last stated 25 Mar 2026 · 6 months ago
Holds David Heinemeier Hansson
Same subject: AI writing
A CEO who sends out an AI-written strategy memo is modelling that it is fine to outsource thinking and strategy.
Last stated 27 Sept 2026 · a week ago
Holds Molly Graham
Same subject: AI writing
A stream of AI-written papers with any error rate at all becomes insufferable, because finding the error costs more than the paper is worth even at ninety-nine percent.
Last stated 30 Jun 2026 · 3 months ago
Holds Grant Sanderson
Same subject: AI writing
AI can now produce in minutes work that used to take weeks to build.
Last stated 26 May 2026 · 4 months ago
Holds Carlos Alexandro Becker
Same subject: open source
A business making hundreds of millions from an open source project owes it money or contribution, and can be refused its trademarks if it gives neither.
Last stated 26 Sept 2024 · 2 years ago
Holds Matt Mullenweg
Same subject: open source
A codec does not really exist until everyone can decode it.
Last stated 31 May 2026 · 4 months ago
Holds Jean-Baptiste Kempf
Same subject: open source
A company that will not accept a project’s terms is free to go and use a more permissive project instead.
Last stated 26 Sept 2024 · 2 years ago
Holds Matt Mullenweg