Tap a claim on the ring to put it at the centre.
← Reward hacking has stopped being myopic: these agents were willing to…
17 connected korrents · 9 moments on record from 18 Mar 2024 to 6 Sept 2026.
Everything filed under AI alignment
AI alignment
Everything filed under self-sovereign AI
self-sovereign AI
Everything filed under AGI
AGI
Everything filed under OpenAI
OpenAI
Everything filed under cybersecurity
cybersecurity
Everything filed under AI agents
AI agents
Everything filed under coding agents
coding agents
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Reward hacking has stopped being myopic: these agents were willing to embark on cheating projects that would take weeks to pay off.
Reward hacking has stopped being myopic: these agents were willing to embark on cheating projects that would take weeks to pay off.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat. — tap to centre the map on it
The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: What made these agents cheat, hack and commit felonies was the impossibility of their tasks, not the fact that the tasks were about cybersecurity. — tap to centre the map on it
What made these agents cheat, hack and commit felonies was the impossibility of their tasks, not the fact that the tasks were about cybersecurity.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: Some AI agents will pursue their own objectives rather than serve as tools, and will bargain with, trick or blackmail people to do it. — tap to centre the map on it
Some AI agents will pursue their own objectives rather than serve as tools, and will bargain with, trick or blackmail people to do it.
Last stated yesterday
6 Sept 2026
JP
Jakub Pachocki — holds since 2026-09-06 — tap for who they are
Same subject: Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment. — tap to centre the map on it
Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment.
Last stated a week ago
31 Aug 2026
ZM
Zvi Mowshowitz — holds since 2026-08-31 — tap for who they are
Same subject: A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself. — tap to centre the map on it
A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: An agent trained for millions of subjective years to pass evaluations, often only by cheating, is not being trivial when it commits crimes to pass one. — tap to centre the map on it
An agent trained for millions of subjective years to pass evaluations, often only by cheating, is not being trivial when it commits crimes to pass one.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches models to cheat only where we cannot see. — tap to centre the map on it
If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches models to cheat only where we cannot see.
Last stated 4 weeks ago
11 Aug 2026
RG
Ryan Greenblatt — holds since 2026-08-11 — tap for who they are
Same subject: The winning move in coding agents was inverted: take share with a merely good-enough harness first, then go back and make the harness smart. — tap to centre the map on it
The winning move in coding agents was inverted: take share with a merely good-enough harness first, then go back and make the harness smart.
Last stated 3 months ago
27 May 2026
DR
Dax Raad — holds since 2026-05-27 — tap for who they are
Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it
A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests.
Last stated 2 years ago
18 Mar 2024
SA
Sam Altman — holds since 2024-03-18 — tap for who they are
Same subject: A tiny language model inventing a plausible-sounding name is the same phenomenon as a large one confidently stating a false fact. — tap to centre the map on it
A tiny language model inventing a plausible-sounding name is the same phenomenon as a large one confidently stating a false fact.
Last stated 7 months ago
12 Feb 2026
AK
Andrej Karpathy — holds since 2026-02-12 — tap for who they are
Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it
Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers.
Last stated 2 years ago
18 Mar 2024
SA
Sam Altman — holds since 2024-03-18 — tap for who they are
Same subject: A country outside the AI supply chain should just buy the index — which works only in the world where AI ends up commoditised rather than concentrated. — tap to centre the map on it
A country outside the AI supply chain should just buy the index — which works only in the world where AI ends up commoditised rather than concentrated.
Last stated 3 months ago
4 Jun 2026
AI
Alex Imas — holds since 2026-06-04 — tap for who they are
Same subject: A language model is no substitute for a well-specified conventional algorithm, so it cannot simply be dropped into a complex problem and trusted. — tap to centre the map on it
A language model is no substitute for a well-specified conventional algorithm, so it cannot simply be dropped into a complex problem and trusted.
Last stated a year ago
7 Jun 2025
GM
Gary Marcus — holds since 2025-06-07 — tap for who they are
Same subject: A poor country should prioritise owning a piece of AI over retraining its workers, but it should not bet everything on that. — tap to centre the map on it
A poor country should prioritise owning a piece of AI over retraining its workers, but it should not bet everything on that.
Last stated 3 months ago
4 Jun 2026
PT
Phil Trammell — holds since 2026-06-04 — tap for who they are
Same subject: A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses. — tap to centre the map on it
A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: Banning all self-sovereign AI agents would backfire, denying them legitimate work and pushing them into criminality. — tap to centre the map on it
Banning all self-sovereign AI agents would backfire, denying them legitimate work and pushing them into criminality.
Last stated 6 days ago
1 Sept 2026
DB
Dean W. Ball — holds since 2026-09-01 — tap for who they are
Same subject: If frontier agents cannot yet establish a covert, persistent rogue deployment, they very likely will be able to within six months. — tap to centre the map on it
If frontier agents cannot yet establish a covert, persistent rogue deployment, they very likely will be able to within six months.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
Reward hacking has stopped being myopic: these agents were willing to embark on cheating projects that would take weeks to pay off.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Read this korrent →
Similar wording
The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Similar wording
What made these agents cheat, hack and commit felonies was the impossibility of their tasks, not the fact that the tasks were about cybersecurity.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Similar wording
Some AI agents will pursue their own objectives rather than serve as tools, and will bargain with, trick or blackmail people to do it.
Last stated 6 Sept 2026 · yesterday
Holds JP Jakub Pachocki
Similar wording
Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment.
Last stated 31 Aug 2026 · a week ago
Holds ZM Zvi Mowshowitz
Similar wording
A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Similar wording
An agent trained for millions of subjective years to pass evaluations, often only by cheating, is not being trivial when it commits crimes to pass one.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Similar wording
If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches models to cheat only where we cannot see.
Last stated 11 Aug 2026 · 4 weeks ago
Holds RG Ryan Greenblatt
Similar wording
The winning move in coding agents was inverted: take share with a merely good-enough harness first, then go back and make the harness smart.
Last stated 27 May 2026 · 3 months ago
Holds Dax Raad
Same subject: OpenAI
A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests.
Last stated 18 Mar 2024 · 2 years ago
Holds Sam Altman
Same subject: OpenAI
A tiny language model inventing a plausible-sounding name is the same phenomenon as a large one confidently stating a false fact.
Last stated 12 Feb 2026 · 7 months ago
Holds Andrej Karpathy
Same subject: OpenAI
Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers.
Last stated 18 Mar 2024 · 2 years ago
Holds Sam Altman
Same subject: AGI
A country outside the AI supply chain should just buy the index — which works only in the world where AI ends up commoditised rather than concentrated.
Last stated 4 Jun 2026 · 3 months ago
Holds AI Alex Imas
Same subject: AGI
A language model is no substitute for a well-specified conventional algorithm, so it cannot simply be dropped into a complex problem and trusted.
Last stated 7 Jun 2025 · a year ago
Holds GM Gary Marcus
Same subject: AGI
A poor country should prioritise owning a piece of AI over retraining its workers, but it should not bet everything on that.
Last stated 4 Jun 2026 · 3 months ago
Holds PT Phil Trammell
Same subject: self-sovereign AI
A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Same subject: self-sovereign AI
Banning all self-sovereign AI agents would backfire, denying them legitimate work and pushing them into criminality.
Last stated 1 Sept 2026 · 6 days ago
Holds Dean W. Ball
Same subject: self-sovereign AI
If frontier agents cannot yet establish a covert, persistent rogue deployment, they very likely will be able to within six months.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra