Tap a claim on the ring to put it at the centre.
← Success is probably measured wrongly when it is judged by the amount of output or the quality of that output.
17 connected korrents · 14 moments from 19 Aug 2004 to 3 Sept 2026.
Everything filed under benchmarks
benchmarks
Everything filed under management
management
Everything filed under ARC-AGI
ARC-AGI
Everything filed under venture capital
venture capital
Everything filed under happiness
happiness
Everything filed under surgical outcomes
surgical outcomes
Everything filed under AI and science
AI and science
Everything filed under education
education
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Success is probably measured wrongly when it is judged by the amount of output or the quality of that output.
Success is probably measured wrongly when it is judged by the amount of output or the quality of that output.
Last stated 2 months ago
21 Jul 2026
PS
Pablo Stanley — holds since 2026-07-21 — tap for who they are
Same subject: Many efforts fail because they do not focus on the right measure or do not invest enough in measuring accurately — tap to centre the map on it
Many efforts fail because they do not focus on the right measure or do not invest enough in measuring accurately
BG
Bill Gates — holds since c. 2013 — tap for who they are
Same subject: Numbers offer a false precision: which metrics you choose to look at, and whether a five percent rise is good, are both art rather than science. — tap to centre the map on it
Numbers offer a false precision: which metrics you choose to look at, and whether a five percent rise is good, are both art rather than science.
Last stated a year ago
21 Sept 2025
JZ
Julie Zhuo — holds since 2025-09-21 — tap for who they are
Same subject: When the data and the anecdotes disagree, the anecdotes are usually right and the metric is usually measuring the wrong thing. — tap to centre the map on it
When the data and the anecdotes disagree, the anecdotes are usually right and the metric is usually measuring the wrong thing.
Last stated 3 years ago
14 Dec 2023
JB
Jeff Bezos — holds since 2023-12-14 — tap for who they are
Same subject: Venture capitalists are judged on the price they entered and exited at, not on the quality of the business they built. — tap to centre the map on it
Venture capitalists are judged on the price they entered and exited at, not on the quality of the business they built.
Last stated a month ago
2 Sept 2026
AD
Aswath Damodaran — holds since 2026-09-02 — tap for who they are
Same subject: People seem to measure by false standards, seeking power, success and riches while undervaluing what is truly precious in life — tap to centre the map on it
People seem to measure by false standards, seeking power, success and riches while undervaluing what is truly precious in life
SF
Sigmund Freud — holds since c. 1930 — tap for who they are
Same subject: Measuring outcomes of medical or surgical interventions is part of good practice, but publishing individual doctors' results remains controversial. — tap to centre the map on it
Measuring outcomes of medical or surgical interventions is part of good practice, but publishing individual doctors' results remains controversial.
Last stated 22 years ago
19 Aug 2004
DS
David Spiegelhalter — holds since 2004-08-19 — tap for who they are
Same subject: Failing to adjust raw data for known measurement biases produces inaccurate results. — tap to centre the map on it
Failing to adjust raw data for known measurement biases produces inaccurate results.
Last stated 2 months ago
27 Jul 2026
AD
Andrew Dessler — holds since 2026-07-27 — tap for who they are
Same subject: A school should be judged on how fast its students are learning, not on how much they already know when measured. — tap to centre the map on it
A school should be judged on how fast its students are learning, not on how much they already know when measured.
Last stated a month ago
31 Aug 2026
JL
Joe Liemandt — holds since 2026-08-31 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated a month ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A carmaker's claim to be the safest is mostly an artefact of comparing a new car against a fleet average twelve years old. — tap to centre the map on it
A carmaker's claim to be the safest is mostly an artefact of comparing a new car against a fleet average twelve years old.
Last stated 3 years ago
19 Dec 2023
PK
Philip Koopman — holds since 2023-12-19 — tap for who they are
Same subject: A model's capability is now a function of how much money you spend on it, so asking what a model can do means nothing until you name a budget. — tap to centre the map on it
A model's capability is now a function of how much money you spend on it, so asking what a model can do means nothing until you name a budget.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores. — tap to centre the map on it
Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores.
Last stated 2 years ago
13 May 2024
SR
Sebastian Ruder — holds since 2024-05-13 — tap for who they are
Same subject: Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used. — tap to centre the map on it
Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans. — tap to centre the map on it
The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: Combining an avocado and a chair into one image is evidence a model conceptually understands both, not that it memorised pictures of them. — tap to centre the map on it
Combining an avocado and a chair into one image is evidence a model conceptually understands both, not that it memorised pictures of them.
Last stated 2 months ago
12 Aug 2026
CF
Chelsea Finn — holds since 2026-08-12 — tap for who they are
Same subject: Passing ARC-AGI does not amount to achieving AGI: o3 still fails on some very easy tasks, indicating fundamental differences from human intelligence. — tap to centre the map on it
Passing ARC-AGI does not amount to achieving AGI: o3 still fails on some very easy tasks, indicating fundamental differences from human intelligence.
Last stated 2 years ago
20 Dec 2024
FC
François Chollet — holds since 2024-12-20 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2004 to today (stretched back to the oldest claim here) — full is today a face: someone who holds the claim — tap it for who they are
At the centre
Success is probably measured wrongly when it is judged by the amount of output or the quality of that output.
Last stated 21 Jul 2026 · 2 months ago
Holds Pablo Stanley
Read this korrent →
Similar wording
Many efforts fail because they do not focus on the right measure or do not invest enough in measuring accurately
Last stated c. 2013
Holds Bill Gates
Similar wording
Numbers offer a false precision: which metrics you choose to look at, and whether a five percent rise is good, are both art rather than science.
Last stated 21 Sept 2025 · a year ago
Holds Julie Zhuo
Similar wording
When the data and the anecdotes disagree, the anecdotes are usually right and the metric is usually measuring the wrong thing.
Last stated 14 Dec 2023 · 3 years ago
Holds Jeff Bezos
Similar wording
Venture capitalists are judged on the price they entered and exited at, not on the quality of the business they built.
Last stated 2 Sept 2026 · a month ago
Holds Aswath Damodaran
Similar wording
People seem to measure by false standards, seeking power, success and riches while undervaluing what is truly precious in life
Last stated c. 1930
Holds Sigmund Freud
Similar wording
Measuring outcomes of medical or surgical interventions is part of good practice, but publishing individual doctors' results remains controversial.
Last stated 19 Aug 2004 · 22 years ago
Holds David Spiegelhalter
Similar wording
Failing to adjust raw data for known measurement biases produces inaccurate results.
Last stated 27 Jul 2026 · 2 months ago
Holds Andrew Dessler
Similar wording
A school should be judged on how fast its students are learning, not on how much they already know when measured.
Last stated 31 Aug 2026 · a month ago
Holds Joe Liemandt
Same subject: benchmarks
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 26 Jun 2026 · 3 months ago
Holds Noam Brown
Same subject: benchmarks
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 Sept 2026 · a month ago
Holds Dax Raad
Same subject: benchmarks
A carmaker's claim to be the safest is mostly an artefact of comparing a new car against a fleet average twelve years old.
Last stated 19 Dec 2023 · 3 years ago
Holds Philip Koopman
Same subject: scaling laws
A model's capability is now a function of how much money you spend on it, so asking what a model can do means nothing until you name a budget.
Last stated 26 Jun 2026 · 3 months ago
Holds Noam Brown
Same subject: scaling laws
Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores.
Last stated 13 May 2024 · 2 years ago
Holds Sebastian Ruder
Same subject: scaling laws
Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used.
Last stated 26 Jun 2026 · 3 months ago
Holds Noam Brown
Same subject: ARC-AGI
The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans.
Last stated 5 Nov 2019 · 7 years ago
Holds François Chollet
Same subject: ARC-AGI
Combining an avocado and a chair into one image is evidence a model conceptually understands both, not that it memorised pictures of them.
Last stated 12 Aug 2026 · 2 months ago
Holds Chelsea Finn
Same subject: ARC-AGI
Passing ARC-AGI does not amount to achieving AGI: o3 still fails on some very easy tasks, indicating fundamental differences from human intelligence.
Last stated 20 Dec 2024 · 2 years ago
Holds François Chollet