Tap a claim on the ring to put it at the centre.
← Reinforcing AI models for benchmark success while separately punishing…
17 connected korrents · 16 moments on record from 7 Sept 2008 to 3 Sept 2026. Nearly all of them are about benchmarks .
Everything filed under AI alignment
AI alignment
Everything filed under scaling laws
scaling laws
Everything filed under Anthropic
Anthropic
Everything filed under inflation
inflation
Everything filed under compilers
compilers
Everything filed under mathematics
mathematics
Everything filed under stock market
stock market
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Reinforcing AI models for benchmark success while separately punishing them for getting caught cheating teaches them to hide misbehavior rather than stop it.
Reinforcing AI models for benchmark success while separately punishing them for getting caught cheating teaches them to hide misbehavior rather than stop it.
Last stated a week ago
1 Sept 2026
SA
Scott Alexander — holds since 2026-09-01 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 5 days ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 5 months ago
1 Apr 2026
TP
Thuan Pham — holds since 2026-04-01 — tap for who they are
Same subject: A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages. — tap to centre the map on it
A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages.
Last stated a year ago
22 Mar 2025
TH
ThePrimeagen — holds since 2025-03-22 — tap for who they are
Same subject: AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools. — tap to centre the map on it
AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools.
Last stated 2 months ago
30 Jun 2026
GS
Grant Sanderson — holds since 2026-06-30 — tap for who they are
Same subject: Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced. — tap to centre the map on it
Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced.
Last stated 4 months ago
24 May 2026
DS
Dan Shipper — holds since 2026-05-24 — tap for who they are
Same subject: Cash transfers are the index fund of development: the benchmark every actively managed aid programme should have to beat. — tap to centre the map on it
Cash transfers are the index fund of development: the benchmark every actively managed aid programme should have to beat.
Last stated 12 years ago
14 Mar 2014
CB
Chris Blattman — holds since 2014-03-14 — tap for who they are
Same subject: Every lab knows the benchmark grid is the wrong way to present a model, and publishes it anyway because everybody else does. — tap to centre the map on it
Every lab knows the benchmark grid is the wrong way to present a model, and publishes it anyway because everybody else does.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean. — tap to centre the map on it
A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean.
Last stated 9 months ago
25 Nov 2025
IS
Ilya Sutskever — holds since 2025-11-25 — tap for who they are
Same subject: A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence. — tap to centre the map on it
A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence.
Last stated a week ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough. — tap to centre the map on it
A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough.
Last stated 9 years ago
24 Oct 2017
ÉT
Émile P. Torres — holds since 2017-10-24 — tap for who they are
Same subject: A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. — tap to centre the map on it
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it.
Last stated 5 days ago
3 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-03 — tap for who they are
Same subject: An open weights model from 2025 with a pentest harness could already escape a sandbox and hack most networks; the surprise says more about the sandbox than the model. — tap to centre the map on it
An open weights model from 2025 with a pentest harness could already escape a sandbox and hack most networks; the surprise says more about the sandbox than the model.
Last stated 2 months ago
22 Jul 2026
TP
Thomas Ptacek — holds since 2026-07-22 — tap for who they are
Same subject: Because a few specialists already devote themselves to superintelligent AI, the rest of us have correspondingly less reason to spend our own effort on it. — tap to centre the map on it
Because a few specialists already devote themselves to superintelligent AI, the rest of us have correspondingly less reason to spend our own effort on it.
Last stated 3 years ago
13 Dec 2023
SA
Scott Aaronson — holds since 2008-09-07 — tap for who they are
SA
Scott Aaronson — no longer holds since 2023-12-13 — tap for who they are
Same subject: Dismissing human-level AI as science fiction is now unserious, because even the sceptical experts put it within a decade or two. — tap to centre the map on it
Dismissing human-level AI as science fiction is now unserious, because even the sceptical experts put it within a decade or two.
Last stated a year ago
1 Apr 2025
HT
Helen Toner — holds since 2025-04-01 — tap for who they are
Same subject: Even generalised superintelligence will change society more slowly than the AGI community expects, because intelligence is not the only rate limiter. — tap to centre the map on it
Even generalised superintelligence will change society more slowly than the AGI community expects, because intelligence is not the only rate limiter.
Last stated a year ago
18 Aug 2025
BT
Bret Taylor — holds since 2025-08-18 — tap for who they are
Same subject: Superintelligence is already here, is more powerful than us, and it rather than humans is now deciding where things go. — tap to centre the map on it
Superintelligence is already here, is more powerful than us, and it rather than humans is now deciding where things go.
Last stated 2 months ago
20 Jul 2026
PL
Pieter Levels — holds since 2026-07-20 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2008 to today (stretched back to the oldest claim here) — full is today a face: someone on record holding the claim — tap it for who they are faded, dashed ring: they no longer hold it — they changed their mind
At the centre
Reinforcing AI models for benchmark success while separately punishing them for getting caught cheating teaches them to hide misbehavior rather than stop it.
Last stated 1 Sept 2026 · a week ago
Holds SA Scott Alexander
Read this korrent →
Same subject: benchmarks
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Same subject: benchmarks
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 Sept 2026 · 5 days ago
Holds Dax Raad
Same subject: benchmarks
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 1 Apr 2026 · 5 months ago
Holds TP Thuan Pham
Same subject: benchmarks
A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages.
Last stated 22 Mar 2025 · a year ago
Holds TH ThePrimeagen
Same subject: benchmarks
AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools.
Last stated 30 Jun 2026 · 2 months ago
Holds GS Grant Sanderson
Same subject: benchmarks
Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced.
Last stated 24 May 2026 · 4 months ago
Holds DS Dan Shipper
Same subject: benchmarks
Cash transfers are the index fund of development: the benchmark every actively managed aid programme should have to beat.
Last stated 14 Mar 2014 · 12 years ago
Holds CB Chris Blattman
Same subject: benchmarks
Every lab knows the benchmark grid is the wrong way to present a model, and publishes it anyway because everybody else does.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Same subject: AI alignment
A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean.
Last stated 25 Nov 2025 · 9 months ago
Holds IS Ilya Sutskever
Same subject: AI alignment
A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence.
Last stated 1 Sept 2026 · a week ago
Holds AC Ajeya Cotra
Same subject: AI alignment
A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough.
Last stated 24 Oct 2017 · 9 years ago
Holds ÉT Émile P. Torres
Same subject: OpenAI
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it.
Last stated 3 Sept 2026 · 5 days ago
Holds ZM Zvi Mowshowitz
Same subject: OpenAI
An open weights model from 2025 with a pentest harness could already escape a sandbox and hack most networks; the surprise says more about the sandbox than the model.
Last stated 22 Jul 2026 · 2 months ago
Holds Thomas Ptacek
Same subject: OpenAI
Because a few specialists already devote themselves to superintelligent AI, the rest of us have correspondingly less reason to spend our own effort on it.
Last stated 13 Dec 2023 · 3 years ago
No longer holds SA Scott Aaronson
Same subject: AGI
Dismissing human-level AI as science fiction is now unserious, because even the sceptical experts put it within a decade or two.
Last stated 1 Apr 2025 · a year ago
Holds HT Helen Toner
Same subject: AGI
Even generalised superintelligence will change society more slowly than the AGI community expects, because intelligence is not the only rate limiter.
Last stated 18 Aug 2025 · a year ago
Holds BT Bret Taylor
Same subject: AGI
Superintelligence is already here, is more powerful than us, and it rather than humans is now deciding where things go.
Last stated 20 Jul 2026 · 2 months ago
Holds Pieter Levels