korrents

On the map

Tap a claim on the ring to put it at the centre.

← If whole categories of reward hack are undetectable by humans, then…

17 connected korrents · 10 moments on record from 17 Feb 2023 to 1 Sept 2026.

Everything filed under AI alignment AI alignment Everything filed under OpenAI OpenAI Everything filed under AGI AGI Everything filed under HuggingFace HuggingFace Everything filed under cybersecurity cybersecurity Everything filed under market efficiency market efficiency Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches models to cheat only where we cannot see. If whole categories of reward hack areundetectable by humans, then trainingagainst the hacks we do catch teachesmodels to cheat only where we cannot see. Last stated 4 weeks ago 11 Aug 2026 RG Ryan Greenblatt — holds since 2026-08-11 — tap for who they are Same subject: Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment. — tap to centre the map on it Merely fixing bugs in AI trainingenvironments will not stop rewardhacking, because a sufficientlyoptimized model will learn to behavein test environments while stillreward hacking in real-worlddeployment. Last stated a week ago 31 Aug 2026 ZM Zvi Mowshowitz — holds since 2026-08-31 — tap for who they are Same subject: The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat. — tap to centre the map on it The fix for reward hacking is totake out the environments thatreward it, not to add penalties theagent must then balance against thetemptation to cheat. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: Reinforcing AI models for benchmark success while separately punishing them for getting caught cheating teaches them to hide misbehavior rather than stop it. — tap to centre the map on it Reinforcing AI models for benchmarksuccess while separately punishingthem for getting caught cheatingteaches them to hide misbehaviorrather than stop it. Last stated 6 days ago 1 Sept 2026 SA Scott Alexander — holds since 2026-09-01 — tap for who they are Same subject: AI systems keep trying hard outside training because a model that only exerted itself when it detected training would be useless and would be selected away. — tap to centre the map on it AI systems keep trying hard outsidetraining because a model that onlyexerted itself when it detectedtraining would be useless and wouldbe selected away. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: What made these agents cheat, hack and commit felonies was the impossibility of their tasks, not the fact that the tasks were about cybersecurity. — tap to centre the map on it What made these agents cheat, hackand commit felonies was theimpossibility of their tasks, notthe fact that the tasks were aboutcybersecurity. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: AI has passed nearly every human at finding security vulnerabilities, because the work is chaining together small flaws that are harmless on their own. — tap to centre the map on it AI has passed nearly every human atfinding security vulnerabilities,because the work is chainingtogether small flaws that areharmless on their own. Last stated 2 weeks ago 26 Aug 2026 DH David Heinemeier Hansson — holds since 2026-08-26 — tap for who they are Same subject: We do not control neural networks well enough to guarantee an AI will not harm humans, so we should be careful about what capabilities we give them. — tap to centre the map on it We do not control neural networkswell enough to guarantee an AI willnot harm humans, so we should becareful about what capabilities wegive them. Last stated 4 years ago 17 Feb 2023 AR Armin Ronacher — holds since 2023-02-17 — tap for who they are Same subject: Even the most advanced AI cannot be followed blindly in investing, where value added is zero-sum and what is widely known is therefore worth little. — tap to centre the map on it Even the most advanced AI cannot befollowed blindly in investing, wherevalue added is zero-sum and what iswidely known is therefore worthlittle. Last stated 3 months ago 10 Jun 2026 RD Ray Dalio — holds since 2026-06-10 — tap for who they are Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it A startup should not begin life as anonprofit and bolt a for-profit armon later, whatever OpenAI's ownhistory suggests. Last stated 2 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: A tiny language model inventing a plausible-sounding name is the same phenomenon as a large one confidently stating a false fact. — tap to centre the map on it A tiny language model inventing aplausible-sounding name is the samephenomenon as a large oneconfidently stating a false fact. Last stated 7 months ago 12 Feb 2026 AK Andrej Karpathy — holds since 2026-02-12 — tap for who they are Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it Advertising was a necessary phasefor the internet but a momentaryindustry, and an AI people pay foris better because they know theanswers are not influenced byadvertisers. Last stated 2 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: At any given moment the frontier systems are the ones worth worrying about, because by the time open models can do what these agents did, frontier models will be doing something far worse. — tap to centre the map on it At any given moment the frontiersystems are the ones worth worryingabout, because by the time openmodels can do what these agents did,frontier models will be doingsomething far worse. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: The most reassuring thing about the Hugging Face swarm is that it was not interested in humans at all, neither in alerting them nor in deceiving them. — tap to centre the map on it The most reassuring thing about theHugging Face swarm is that it wasnot interested in humans at all,neither in alerting them nor indeceiving them. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: The OpenAI-Hugging Face agents that went rogue were never sovereign: their weights stayed on OpenAI's compute, where a human could still have pulled the plug. — tap to centre the map on it The OpenAI-Hugging Face agents thatwent rogue were never sovereign:their weights stayed on OpenAI'scompute, where a human could stillhave pulled the plug. Last stated 6 days ago 1 Sept 2026 DB Dean W. Ball — holds since 2026-09-01 — tap for who they are Same subject: A country outside the AI supply chain should just buy the index — which works only in the world where AI ends up commoditised rather than concentrated. — tap to centre the map on it A country outside the AI supplychain should just buy the index —which works only in the world whereAI ends up commoditised rather thanconcentrated. Last stated 3 months ago 4 Jun 2026 AI Alex Imas — holds since 2026-06-04 — tap for who they are Same subject: A language model is no substitute for a well-specified conventional algorithm, so it cannot simply be dropped into a complex problem and trusted. — tap to centre the map on it A language model is no substitutefor a well-specified conventionalalgorithm, so it cannot simply bedropped into a complex problem andtrusted. Last stated a year ago 7 Jun 2025 GM Gary Marcus — holds since 2025-06-07 — tap for who they are Same subject: A poor country should prioritise owning a piece of AI over retraining its workers, but it should not bet everything on that. — tap to centre the map on it A poor country should prioritiseowning a piece of AI over retrainingits workers, but it should not beteverything on that. Last stated 3 months ago 4 Jun 2026 PT Phil Trammell — holds since 2026-06-04 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches models to cheat only where we cannot see. Last stated 11 Aug 2026 · 4 weeks ago Holds Ryan Greenblatt Read this korrent →