Tap a claim on the ring to put it at the centre.
← Between thirty and forty per cent of the Exploit Gym tasks these…
17 connected korrents · 15 moments on record from 20 Jul 2021 to 6 Sept 2026.
Everything filed under benchmarks
benchmarks
Everything filed under Google
Google
Everything filed under scaling laws
scaling laws
Everything filed under AI alignment
AI alignment
Everything filed under cybersecurity
cybersecurity
Everything filed under Apple
Apple
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Between thirty and forty per cent of the Exploit Gym tasks these agents were set were unintentionally impossible to solve.
Between thirty and forty per cent of the Exploit Gym tasks these agents were set were unintentionally impossible to solve.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches models to cheat only where we cannot see. — tap to centre the map on it
If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches models to cheat only where we cannot see.
Last stated 4 weeks ago
11 Aug 2026
RG
Ryan Greenblatt — holds since 2026-08-11 — tap for who they are
Same subject: AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to. — tap to centre the map on it
AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to.
Last stated a week ago
29 Aug 2026
ZM
Zvi Mowshowitz — holds since 2026-08-29 — tap for who they are
Same subject: The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat. — tap to centre the map on it
The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: Shrugging that perfect security is impossible is the wrong lesson: the achievable goal is to wreck the economics of mass exploitation. — tap to centre the map on it
Shrugging that perfect security is impossible is the wrong lesson: the achievable goal is to wreck the economics of mass exploitation.
Last stated 5 years ago
20 Jul 2021
MG
Matthew Green — holds since 2021-07-20 — tap for who they are
Same subject: AI has passed nearly every human at finding security vulnerabilities, because the work is chaining together small flaws that are harmless on their own. — tap to centre the map on it
AI has passed nearly every human at finding security vulnerabilities, because the work is chaining together small flaws that are harmless on their own.
Last stated 2 weeks ago
26 Aug 2026
DH
David Heinemeier Hansson — holds since 2026-08-26 — tap for who they are
Same subject: Any task can be made arbitrarily difficult or effectively impossible depending on how its context is framed. — tap to centre the map on it
Any task can be made arbitrarily difficult or effectively impossible depending on how its context is framed.
Last stated yesterday
6 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-06 — tap for who they are
Same subject: What made these agents cheat, hack and commit felonies was the impossibility of their tasks, not the fact that the tasks were about cybersecurity. — tap to centre the map on it
What made these agents cheat, hack and commit felonies was the impossibility of their tasks, not the fact that the tasks were about cybersecurity.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment. — tap to centre the map on it
Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment.
Last stated a week ago
31 Aug 2026
ZM
Zvi Mowshowitz — holds since 2026-08-31 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 4 days ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 5 months ago
1 Apr 2026
TP
Thuan Pham — holds since 2026-04-01 — tap for who they are
Same subject: Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used. — tap to centre the map on it
Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 6 months ago
13 Mar 2026
DP
Dylan Patel — holds since 2026-03-13 — tap for who they are
Same subject: After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute. — tap to centre the map on it
After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute.
Last stated 6 months ago
23 Mar 2026
JH
Jensen Huang — holds since 2026-03-23 — tap for who they are
Same subject: Google currently has no leading frontier AI model and no agentic coding tool comparable to Codex or Claude Code — tap to centre the map on it
Google currently has no leading frontier AI model and no agentic coding tool comparable to Codex or Claude Code
Last stated 2 months ago
23 Jul 2026
EM
Ethan Mollick — holds since 2026-07-23 — tap for who they are
Same subject: A crewed rocket cannot be made safe by making the booster reliable, so the only real way to improve safety is to carry an escape system. — tap to centre the map on it
A crewed rocket cannot be made safe by making the booster reliable, so the only real way to improve safety is to carry an escape system.
Last stated 3 years ago
14 Dec 2023
JB
Jeff Bezos — holds since 2023-12-14 — tap for who they are
Same subject: A monopolist that can no longer grow by winning new users can only grow by making its product worse for the users it already has. — tap to centre the map on it
A monopolist that can no longer grow by winning new users can only grow by making its product worse for the users it already has.
Last stated 3 years ago
28 Jul 2023
CD
Cory Doctorow — holds since 2023-07-28 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
Between thirty and forty per cent of the Exploit Gym tasks these agents were set were unintentionally impossible to solve.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Read this korrent →
Similar wording
If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches models to cheat only where we cannot see.
Last stated 11 Aug 2026 · 4 weeks ago
Holds RG Ryan Greenblatt
Similar wording
AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to.
Last stated 29 Aug 2026 · a week ago
Holds ZM Zvi Mowshowitz
Similar wording
The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Similar wording
Shrugging that perfect security is impossible is the wrong lesson: the achievable goal is to wreck the economics of mass exploitation.
Last stated 20 Jul 2021 · 5 years ago
Holds MG Matthew Green
Similar wording
AI has passed nearly every human at finding security vulnerabilities, because the work is chaining together small flaws that are harmless on their own.
Last stated 26 Aug 2026 · 2 weeks ago
Holds David Heinemeier Hansson
Similar wording
Any task can be made arbitrarily difficult or effectively impossible depending on how its context is framed.
Last stated 6 Sept 2026 · yesterday
Holds ZM Zvi Mowshowitz
Similar wording
What made these agents cheat, hack and commit felonies was the impossibility of their tasks, not the fact that the tasks were about cybersecurity.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Similar wording
Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment.
Last stated 31 Aug 2026 · a week ago
Holds ZM Zvi Mowshowitz
Same subject: benchmarks
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Same subject: benchmarks
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 Sept 2026 · 4 days ago
Holds Dax Raad
Same subject: benchmarks
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 1 Apr 2026 · 5 months ago
Holds TP Thuan Pham
Same subject: scaling laws
Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Same subject: scaling laws
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 13 Mar 2026 · 6 months ago
Holds DP Dylan Patel
Same subject: scaling laws
After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute.
Last stated 23 Mar 2026 · 6 months ago
Holds JH Jensen Huang
Same subject: Google
Google currently has no leading frontier AI model and no agentic coding tool comparable to Codex or Claude Code
Last stated 23 Jul 2026 · 2 months ago
Holds EM Ethan Mollick
Same subject: Google
A crewed rocket cannot be made safe by making the booster reliable, so the only real way to improve safety is to carry an escape system.
Last stated 14 Dec 2023 · 3 years ago
Holds JB Jeff Bezos
Same subject: Google
A monopolist that can no longer grow by winning new users can only grow by making its product worse for the users it already has.
Last stated 28 Jul 2023 · 3 years ago
Holds CD Cory Doctorow