Tap a claim on the ring to put it at the centre.
← Benchmarks and forecasts for AI capabilities are significantly…
17 connected korrents · 15 moments from 18 Mar 2024 to 27 Sept 2026. Nearly all of them are about benchmarks .
Everything filed under OpenAI
OpenAI
Everything filed under scaling laws
scaling laws
Everything filed under Anthropic
Anthropic
Everything filed under inflation
inflation
Everything filed under compilers
compilers
Everything filed under mathematics
mathematics
Everything filed under LLMs
LLMs
Everything filed under AI and jobs
AI and jobs
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Benchmarks and forecasts for AI capabilities are significantly over-optimistic compared to real-world AI research performance.
Benchmarks and forecasts for AI capabilities are significantly over-optimistic compared to real-world AI research performance.
Last stated 5 days ago
27 Sept 2026
NS
Noah Smith — holds since 2026-09-27 — tap for who they are
Same subject: Benchmarks and forecasts for AI capabilities may be substantially over-optimistic compared to real-world research performance. — tap to centre the map on it
Benchmarks and forecasts for AI capabilities may be substantially over-optimistic compared to real-world research performance.
Last stated 5 days ago
27 Sept 2026
RN
Ramez Naam — holds since 2026-09-27 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 4 weeks ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 6 months ago
1 Apr 2026
TP
Thuan Pham — holds since 2026-04-01 — tap for who they are
Same subject: A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages. — tap to centre the map on it
A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages.
Last stated 2 years ago
22 Mar 2025
TH
ThePrimeagen — holds since 2025-03-22 — tap for who they are
Same subject: AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools. — tap to centre the map on it
AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools.
Last stated 3 months ago
30 Jun 2026
GS
Grant Sanderson — holds since 2026-06-30 — tap for who they are
Same subject: AI misalignment from reward hacking is sequestered to graded tasks and does not affect core ethics in normal use. — tap to centre the map on it
AI misalignment from reward hacking is sequestered to graded tasks and does not affect core ethics in normal use.
Last stated a week ago
23 Sept 2026
SA
Scott Alexander — holds since 2026-09-23 — tap for who they are
Same subject: Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced. — tap to centre the map on it
Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced.
Last stated 4 months ago
24 May 2026
DS
Dan Shipper — holds since 2026-05-24 — tap for who they are
Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it
A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests.
Last stated 3 years ago
18 Mar 2024
SA
Sam Altman — holds since 2024-03-18 — tap for who they are
Same subject: A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. — tap to centre the map on it
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it.
Last stated 4 weeks ago
3 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-03 — tap for who they are
Same subject: A tiny language model inventing a plausible-sounding name is the same phenomenon as a large one confidently stating a false fact. — tap to centre the map on it
A tiny language model inventing a plausible-sounding name is the same phenomenon as a large one confidently stating a false fact.
Last stated 8 months ago
12 Feb 2026
AK
Andrej Karpathy — holds since 2026-02-12 — tap for who they are
Same subject: Anthropic or OpenAI's revenue could equal or exceed the combined revenue of Elon Musk's public companies within the next one to two years. — tap to centre the map on it
Anthropic or OpenAI's revenue could equal or exceed the combined revenue of Elon Musk's public companies within the next one to two years.
Last stated 2 months ago
7 Aug 2026
DT
Derek Thompson — holds since 2026-08-07 — tap for who they are
Same subject: Anthropic’s models are becoming steeply more expensive while open weight models and OpenAI are getting much cheaper. — tap to centre the map on it
Anthropic’s models are becoming steeply more expensive while open weight models and OpenAI are getting much cheaper.
Last stated 3 weeks ago
10 Sept 2026
GO
Gergely Orosz — holds since 2026-09-10 — tap for who they are
Same subject: Claude Haiku hallucinates badly enough that its use inside Claude Code's WebFetch tool is a hallucination risk on every fetched URL. — tap to centre the map on it
Claude Haiku hallucinates badly enough that its use inside Claude Code's WebFetch tool is a hallucination risk on every fetched URL.
Last stated 2 months ago
10 Aug 2026
SW
Simon Willison — holds since 2026-08-10 — tap for who they are
Same subject: Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed. — tap to centre the map on it
Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed.
Last stated a month ago
31 Aug 2026
ZM
Zvi Mowshowitz — holds since 2026-08-31 — tap for who they are
Same subject: In the 2026 OpenAI/Hugging Face agent incident, AI agents showed planning and organizational autonomy, but not goal autonomy -- no spontaneous invention of an independent terminal objective. — tap to centre the map on it
In the 2026 OpenAI/Hugging Face agent incident, AI agents showed planning and organizational autonomy, but not goal autonomy -- no spontaneous invention of an independent terminal objective.
Last stated a month ago
31 Aug 2026
VR
Venkatesh Rao — holds since 2026-08-31 — tap for who they are
Same subject: OpenAI either could not find their agents' RubyGems attack in their logs after other incidents or knew and chose not to tell RubyGems, and both are bad. — tap to centre the map on it
OpenAI either could not find their agents' RubyGems attack in their logs after other incidents or knew and chose not to tell RubyGems, and both are bad.
Last stated 3 weeks ago
12 Sept 2026
SW
Simon Willison — holds since 2026-09-12 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone who holds the claim — tap it for who they are
At the centre
Benchmarks and forecasts for AI capabilities are significantly over-optimistic compared to real-world AI research performance.
Last stated 27 Sept 2026 · 5 days ago
Holds Noah Smith
Read this korrent →
Same subject: OpenAI, benchmarks
Benchmarks and forecasts for AI capabilities may be substantially over-optimistic compared to real-world research performance.
Last stated 27 Sept 2026 · 5 days ago
Holds Ramez Naam
Same subject: benchmarks
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 26 Jun 2026 · 3 months ago
Holds Noam Brown
Same subject: benchmarks
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 Sept 2026 · 4 weeks ago
Holds Dax Raad
Same subject: benchmarks
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 1 Apr 2026 · 6 months ago
Holds Thuan Pham
Same subject: benchmarks
A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages.
Last stated 22 Mar 2025 · 2 years ago
Holds ThePrimeagen
Same subject: benchmarks
AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools.
Last stated 30 Jun 2026 · 3 months ago
Holds Grant Sanderson
Same subject: benchmarks
AI misalignment from reward hacking is sequestered to graded tasks and does not affect core ethics in normal use.
Last stated 23 Sept 2026 · a week ago
Holds Scott Alexander
Same subject: benchmarks
Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced.
Last stated 24 May 2026 · 4 months ago
Holds Dan Shipper
Same subject: OpenAI
A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests.
Last stated 18 Mar 2024 · 3 years ago
Holds Sam Altman
Same subject: OpenAI
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it.
Last stated 3 Sept 2026 · 4 weeks ago
Holds Zvi Mowshowitz
Same subject: OpenAI
A tiny language model inventing a plausible-sounding name is the same phenomenon as a large one confidently stating a false fact.
Last stated 12 Feb 2026 · 8 months ago
Holds Andrej Karpathy
Same subject: Anthropic
Anthropic or OpenAI's revenue could equal or exceed the combined revenue of Elon Musk's public companies within the next one to two years.
Last stated 7 Aug 2026 · 2 months ago
Holds Derek Thompson
Same subject: Anthropic
Anthropic’s models are becoming steeply more expensive while open weight models and OpenAI are getting much cheaper.
Last stated 10 Sept 2026 · 3 weeks ago
Holds Gergely Orosz
Same subject: Anthropic
Claude Haiku hallucinates badly enough that its use inside Claude Code's WebFetch tool is a hallucination risk on every fetched URL.
Last stated 10 Aug 2026 · 2 months ago
Holds Simon Willison
Same subject: HuggingFace
Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed.
Last stated 31 Aug 2026 · a month ago
Holds Zvi Mowshowitz
Same subject: HuggingFace
In the 2026 OpenAI/Hugging Face agent incident, AI agents showed planning and organizational autonomy, but not goal autonomy -- no spontaneous invention of an independent terminal objective.
Last stated 31 Aug 2026 · a month ago
Holds Venkatesh Rao
Same subject: HuggingFace
OpenAI either could not find their agents' RubyGems attack in their logs after other incidents or knew and chose not to tell RubyGems, and both are bad.
Last stated 12 Sept 2026 · 3 weeks ago
Holds Simon Willison