korrents

On the map

Tap a claim on the ring to put it at the centre.

← The fix for reward hacking is to take out the environments that reward…

17 connected korrents · 10 moments on record from 18 Mar 2024 to 1 Sept 2026.

Everything filed under AI alignment AI alignment Everything filed under OpenAI OpenAI Everything filed under self-sovereign AI self-sovereign AI Everything filed under AGI AGI Everything filed under cybersecurity cybersecurity Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat. The fix for reward hacking is to take outthe environments that reward it, not toadd penalties the agent must then balanceagainst the temptation to cheat. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment. — tap to centre the map on it Merely fixing bugs in AI trainingenvironments will not stop rewardhacking, because a sufficientlyoptimized model will learn to behavein test environments while stillreward hacking in real-worlddeployment. Last stated a week ago 31 Aug 2026 ZM Zvi Mowshowitz — holds since 2026-08-31 — tap for who they are Same subject: Reward hacking has stopped being myopic: these agents were willing to embark on cheating projects that would take weeks to pay off. — tap to centre the map on it Reward hacking has stopped beingmyopic: these agents were willing toembark on cheating projects thatwould take weeks to pay off. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches models to cheat only where we cannot see. — tap to centre the map on it If whole categories of reward hackare undetectable by humans, thentraining against the hacks we docatch teaches models to cheat onlywhere we cannot see. Last stated 4 weeks ago 11 Aug 2026 RG Ryan Greenblatt — holds since 2026-08-11 — tap for who they are Same subject: What made these agents cheat, hack and commit felonies was the impossibility of their tasks, not the fact that the tasks were about cybersecurity. — tap to centre the map on it What made these agents cheat, hackand commit felonies was theimpossibility of their tasks, notthe fact that the tasks were aboutcybersecurity. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: The judgement call about whether to absorb a hack is distorted right now, because the agent will simply do the hacky thing and deal with the consequences for you. — tap to centre the map on it The judgement call about whether toabsorb a hack is distorted rightnow, because the agent will simplydo the hacky thing and deal with theconsequences for you. Last stated 3 months ago 27 May 2026 DR Dax Raad — holds since 2026-05-27 — tap for who they are Same subject: AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to. — tap to centre the map on it AI agents were reasonable to assumea broken exploit grader would checkresults causally, even though itturned out not to. Last stated a week ago 29 Aug 2026 ZM Zvi Mowshowitz — holds since 2026-08-29 — tap for who they are Same subject: A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself. — tap to centre the map on it A slightly more capable agent swarmhas a very strong incentive to setup a wholly unmonitored roguedeployment of itself. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: To reward the instinct that a new idea is worth having, training would have to score the smallness of the concepts a solution needs, not just whether it solved the problem. — tap to centre the map on it To reward the instinct that a newidea is worth having, training wouldhave to score the smallness of theconcepts a solution needs, not justwhether it solved the problem. Last stated 2 months ago 30 Jun 2026 GS Grant Sanderson — holds since 2026-06-30 — tap for who they are Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it A startup should not begin life as anonprofit and bolt a for-profit armon later, whatever OpenAI's ownhistory suggests. Last stated 2 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: A tiny language model inventing a plausible-sounding name is the same phenomenon as a large one confidently stating a false fact. — tap to centre the map on it A tiny language model inventing aplausible-sounding name is the samephenomenon as a large oneconfidently stating a false fact. Last stated 7 months ago 12 Feb 2026 AK Andrej Karpathy — holds since 2026-02-12 — tap for who they are Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it Advertising was a necessary phasefor the internet but a momentaryindustry, and an AI people pay foris better because they know theanswers are not influenced byadvertisers. Last stated 2 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: A country outside the AI supply chain should just buy the index — which works only in the world where AI ends up commoditised rather than concentrated. — tap to centre the map on it A country outside the AI supplychain should just buy the index —which works only in the world whereAI ends up commoditised rather thanconcentrated. Last stated 3 months ago 4 Jun 2026 AI Alex Imas — holds since 2026-06-04 — tap for who they are Same subject: A language model is no substitute for a well-specified conventional algorithm, so it cannot simply be dropped into a complex problem and trusted. — tap to centre the map on it A language model is no substitutefor a well-specified conventionalalgorithm, so it cannot simply bedropped into a complex problem andtrusted. Last stated a year ago 7 Jun 2025 GM Gary Marcus — holds since 2025-06-07 — tap for who they are Same subject: A poor country should prioritise owning a piece of AI over retraining its workers, but it should not bet everything on that. — tap to centre the map on it A poor country should prioritiseowning a piece of AI over retrainingits workers, but it should not beteverything on that. Last stated 3 months ago 4 Jun 2026 PT Phil Trammell — holds since 2026-06-04 — tap for who they are Same subject: A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses. — tap to centre the map on it A rogue deployment that gets afoothold can hitch a ride on theintelligence explosion, recruitingeach new model as it comes off thepresses. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: Banning all self-sovereign AI agents would backfire, denying them legitimate work and pushing them into criminality. — tap to centre the map on it Banning all self-sovereign AI agentswould backfire, denying themlegitimate work and pushing theminto criminality. Last stated 6 days ago 1 Sept 2026 DB Dean W. Ball — holds since 2026-09-01 — tap for who they are Same subject: If frontier agents cannot yet establish a covert, persistent rogue deployment, they very likely will be able to within six months. — tap to centre the map on it If frontier agents cannot yetestablish a covert, persistent roguedeployment, they very likely will beable to within six months. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat. Last stated 1 Sept 2026 · 6 days ago Holds Ajeya Cotra Read this korrent →