Tap a claim on the ring to put it at the centre.
← The tit-for-tat strategy, despite its reputation as the most…
12 connected korrents · 11 moments from 5 Nov 2019 to 3 Sept 2026.
Everything filed under benchmarks
benchmarks
Everything filed under ARC-AGI
ARC-AGI
Everything filed under mental models
mental models
Everything filed under design
design
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: The tit-for-tat strategy, despite its reputation as the most successful approach in game-theory strategy comparisons, ranks poorly when tested against the full space of possible strategies.
The tit-for-tat strategy, despite its reputation as the most successful approach in game-theory strategy comparisons, ranks poorly when tested against the full space of possible strategies.
Last stated 4 months ago
4 Jun 2026
SW
Stephen Wolfram — holds since 2026-06-04 — tap for who they are
Same subject: Tit-for-tat in the iterated prisoner's dilemma is the piece of game theory most worth knowing. — tap to centre the map on it
Tit-for-tat in the iterated prisoner's dilemma is the piece of game theory most worth knowing.
Last stated 7 years ago
28 Dec 2019
NR
Naval Ravikant — holds since 2019-12-28 — tap for who they are
Same subject: A board game conquers the world when chance and strategy are in balance in it. — tap to centre the map on it
A board game conquers the world when chance and strategy are in balance in it.
Last stated 10 months ago
12 Dec 2025
IF
Irving Finkel — holds since 2025-12-12 — tap for who they are
Same subject: Pursuing a very low success-rate strategy can still be rational if a rare win is valuable enough. — tap to centre the map on it
Pursuing a very low success-rate strategy can still be rational if a rare win is valuable enough.
Last stated 5 years ago
22 Feb 2022
NN
Neel Nanda — holds since 2022-02-22 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated a month ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A carmaker's claim to be the safest is mostly an artefact of comparing a new car against a fleet average twelve years old. — tap to centre the map on it
A carmaker's claim to be the safest is mostly an artefact of comparing a new car against a fleet average twelve years old.
Last stated 3 years ago
19 Dec 2023
PK
Philip Koopman — holds since 2023-12-19 — tap for who they are
Same subject: A model's capability is now a function of how much money you spend on it, so asking what a model can do means nothing until you name a budget. — tap to centre the map on it
A model's capability is now a function of how much money you spend on it, so asking what a model can do means nothing until you name a budget.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores. — tap to centre the map on it
Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores.
Last stated 2 years ago
13 May 2024
SR
Sebastian Ruder — holds since 2024-05-13 — tap for who they are
Same subject: Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used. — tap to centre the map on it
Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans. — tap to centre the map on it
The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: Combining an avocado and a chair into one image is evidence a model conceptually understands both, not that it memorised pictures of them. — tap to centre the map on it
Combining an avocado and a chair into one image is evidence a model conceptually understands both, not that it memorised pictures of them.
Last stated 2 months ago
12 Aug 2026
CF
Chelsea Finn — holds since 2026-08-12 — tap for who they are
Same subject: Passing ARC-AGI does not amount to achieving AGI: o3 still fails on some very easy tasks, indicating fundamental differences from human intelligence. — tap to centre the map on it
Passing ARC-AGI does not amount to achieving AGI: o3 still fails on some very easy tasks, indicating fundamental differences from human intelligence.
Last stated 2 years ago
20 Dec 2024
FC
François Chollet — holds since 2024-12-20 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone who holds the claim — tap it for who they are