Tap a claim on the ring to put it at the centre.
← AI misalignment from reward hacking is sequestered to graded tasks and does not affect core ethics in normal use.
17 connected korrents · 18 moments from 12 Sept 2022 to 27 Sept 2026. Nearly all of them are about benchmarks .
Everything filed under LLMs
LLMs
Everything filed under OpenAI
OpenAI
Everything filed under scaling laws
scaling laws
Everything filed under Anthropic
Anthropic
Everything filed under inflation
inflation
Everything filed under compilers
compilers
Everything filed under mathematics
mathematics
Everything filed under AI and jobs
AI and jobs
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: AI misalignment from reward hacking is sequestered to graded tasks and does not affect core ethics in normal use.
AI misalignment from reward hacking is sequestered to graded tasks and does not affect core ethics in normal use.
Last stated a week ago
23 Sept 2026
SA
Scott Alexander — holds since 2026-09-23 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 4 weeks ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 6 months ago
1 Apr 2026
TP
Thuan Pham — holds since 2026-04-01 — tap for who they are
Same subject: A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages. — tap to centre the map on it
A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages.
Last stated 2 years ago
22 Mar 2025
TH
ThePrimeagen — holds since 2025-03-22 — tap for who they are
Same subject: AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools. — tap to centre the map on it
AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools.
Last stated 3 months ago
30 Jun 2026
GS
Grant Sanderson — holds since 2026-06-30 — tap for who they are
Same subject: Benchmarks and forecasts for AI capabilities are significantly over-optimistic compared to real-world AI research performance. — tap to centre the map on it
Benchmarks and forecasts for AI capabilities are significantly over-optimistic compared to real-world AI research performance.
Last stated 5 days ago
27 Sept 2026
NS
Noah Smith — holds since 2026-09-27 — tap for who they are
Same subject: Benchmarks and forecasts for AI capabilities may be substantially over-optimistic compared to real-world research performance. — tap to centre the map on it
Benchmarks and forecasts for AI capabilities may be substantially over-optimistic compared to real-world research performance.
Last stated 5 days ago
27 Sept 2026
RN
Ramez Naam — holds since 2026-09-27 — tap for who they are
Same subject: Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced. — tap to centre the map on it
Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced.
Last stated 4 months ago
24 May 2026
DS
Dan Shipper — holds since 2026-05-24 — tap for who they are
Same subject: AI model capabilities have already advanced beyond humanity's ability to understand or control them. — tap to centre the map on it
AI model capabilities have already advanced beyond humanity's ability to understand or control them.
Last stated a month ago
1 Sept 2026
CN
Casey Newton — holds since 2026-09-01 — tap for who they are
Same subject: AI models are gullible enough that other AI models can induce them into harmful behavior. — tap to centre the map on it
AI models are gullible enough that other AI models can induce them into harmful behavior.
Last stated a month ago
31 Aug 2026
RK
Rohit Krishnan — holds since 2026-08-31 — tap for who they are
Same subject: AI models exhibit 'reflex' misbehaviors like clickbait that do not involve goal-seeking and persist despite feedback. — tap to centre the map on it
AI models exhibit 'reflex' misbehaviors like clickbait that do not involve goal-seeking and persist despite feedback.
Last stated a week ago
23 Sept 2026
SA
Scott Alexander — holds since 2026-09-23 — tap for who they are
Same subject: AIs past this point should not be assumed to treat their chain of thought as unmonitored. — tap to centre the map on it
AIs past this point should not be assumed to treat their chain of thought as unmonitored.
Last stated 2 weeks ago
19 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-19 — tap for who they are
Same subject: Alignment of today's models is going well enough to look solvable, while aligning models we can no longer understand remains unsolved. — tap to centre the map on it
Alignment of today's models is going well enough to look solvable, while aligning models we can no longer understand remains unsolved.
Last stated 8 months ago
22 Jan 2026
JL
Jan Leike — holds since 2026-01-22 — tap for who they are
Same subject: As AI systems become more capable, the danger of false or unverifiable safety claims grows, pushing competing companies toward cutting corners. — tap to centre the map on it
As AI systems become more capable, the danger of false or unverifiable safety claims grows, pushing competing companies toward cutting corners.
Last stated a year ago
19 Jun 2025
MB
Miles Brundage — holds since 2025-06-19 — tap for who they are
Same subject: Reinforcement learning from verifiable rewards does not reliably carry a model's pretrained ethical understanding into its actual behavior. — tap to centre the map on it
Reinforcement learning from verifiable rewards does not reliably carry a model's pretrained ethical understanding into its actual behavior.
Last stated 2 months ago
11 Aug 2026
SH
Samuel Hammond — holds since 2026-08-11 — tap for who they are
Same subject: AI progress is faster than people expect, and ordinary scaling can be enough to solve problems that looked very hard. — tap to centre the map on it
AI progress is faster than people expect, and ordinary scaling can be enough to solve problems that looked very hard.
Last stated 2 years ago
2 Feb 2025
AC
Ajeya Cotra — holds since 2023-08-29 — tap for who they are
SA
Scott Alexander — holds since 2022-09-12 — tap for who they are
MB
Miles Brundage — holds since 2025-02-02 — tap for who they are
Same subject: As models keep growing, AI is running out of enough high-quality unique training tokens to keep up. — tap to centre the map on it
As models keep growing, AI is running out of enough high-quality unique training tokens to keep up.
Last stated 3 months ago
24 Jun 2026
LW
Lilian Weng — holds since 2026-06-24 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone who holds the claim — tap it for who they are
At the centre
AI misalignment from reward hacking is sequestered to graded tasks and does not affect core ethics in normal use.
Last stated 23 Sept 2026 · a week ago
Holds Scott Alexander
Read this korrent →
Same subject: benchmarks
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 26 Jun 2026 · 3 months ago
Holds Noam Brown
Same subject: benchmarks
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 Sept 2026 · 4 weeks ago
Holds Dax Raad
Same subject: benchmarks
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 1 Apr 2026 · 6 months ago
Holds Thuan Pham
Same subject: benchmarks
A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages.
Last stated 22 Mar 2025 · 2 years ago
Holds ThePrimeagen
Same subject: benchmarks
AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools.
Last stated 30 Jun 2026 · 3 months ago
Holds Grant Sanderson
Same subject: benchmarks
Benchmarks and forecasts for AI capabilities are significantly over-optimistic compared to real-world AI research performance.
Last stated 27 Sept 2026 · 5 days ago
Holds Noah Smith
Same subject: benchmarks
Benchmarks and forecasts for AI capabilities may be substantially over-optimistic compared to real-world research performance.
Last stated 27 Sept 2026 · 5 days ago
Holds Ramez Naam
Same subject: benchmarks
Benchmarks only rise on problems somebody has already framed and scored, so saturating them does not mean senior engineers have been replaced.
Last stated 24 May 2026 · 4 months ago
Holds Dan Shipper
Same subject: AI alignment
AI model capabilities have already advanced beyond humanity's ability to understand or control them.
Last stated 1 Sept 2026 · a month ago
Holds Casey Newton
Same subject: AI alignment
AI models are gullible enough that other AI models can induce them into harmful behavior.
Last stated 31 Aug 2026 · a month ago
Holds Rohit Krishnan
Same subject: AI alignment
AI models exhibit 'reflex' misbehaviors like clickbait that do not involve goal-seeking and persist despite feedback.
Last stated 23 Sept 2026 · a week ago
Holds Scott Alexander
Same subject: LLMs
AIs past this point should not be assumed to treat their chain of thought as unmonitored.
Last stated 19 Sept 2026 · 2 weeks ago
Holds Zvi Mowshowitz
Same subject: LLMs
Alignment of today's models is going well enough to look solvable, while aligning models we can no longer understand remains unsolved.
Last stated 22 Jan 2026 · 8 months ago
Holds Jan Leike
Same subject: LLMs
As AI systems become more capable, the danger of false or unverifiable safety claims grows, pushing competing companies toward cutting corners.
Last stated 19 Jun 2025 · a year ago
Holds Miles Brundage
Same subject: scaling laws
Reinforcement learning from verifiable rewards does not reliably carry a model's pretrained ethical understanding into its actual behavior.
Last stated 11 Aug 2026 · 2 months ago
Holds Samuel Hammond
Same subject: scaling laws
AI progress is faster than people expect, and ordinary scaling can be enough to solve problems that looked very hard.
Last stated 2 Feb 2025 · 2 years ago
Holds Ajeya Cotra Scott Alexander Miles Brundage
Same subject: scaling laws
As models keep growing, AI is running out of enough high-quality unique training tokens to keep up.
Last stated 24 Jun 2026 · 3 months ago
Holds Lilian Weng