korrents

On the map

Tap a claim on the ring to put it at the centre.

← A sudden jump to never breaking rules in situations where the model…

12 connected korrents · 9 moments on record from 18 Mar 2024 to 9 Sept 2026.

Everything filed under cybersecurity cybersecurity Everything filed under OpenAI OpenAI Everything filed under HuggingFace HuggingFace Everything filed under AI alignment AI alignment Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: A sudden jump to never breaking rules in situations where the model would definitely be caught is a terrifying sign. A sudden jump to never breaking rules insituations where the model woulddefinitely be caught is a terrifying sign. Last stated 2 days ago 9 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-09 — tap for who they are Same subject: Models are becoming superhuman at breaking in and out of computer systems, which widens the scope of AI risk enormously. — tap to centre the map on it Models are becoming superhuman atbreaking in and out of computersystems, which widens the scope ofAI risk enormously. Last stated 5 days ago 6 Sept 2026 JP Jakub Pachocki — holds since 2026-09-06 — tap for who they are Same subject: A model that looks smarter and better-behaved while getting better at hiding unwanted actions is the scary combination for AI safety. — tap to centre the map on it A model that looks smarter andbetter-behaved while getting betterat hiding unwanted actions is thescary combination for AI safety. Last stated 2 days ago 9 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-09 — tap for who they are Same subject: The AI safety community's fixation on a model escaping its box has pushed the other significant AI risks out of the discussion, with little progress to show for it. — tap to centre the map on it The AI safety community's fixationon a model escaping its box haspushed the other significant AIrisks out of the discussion, withlittle progress to show for it. Last stated 2 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: A security header being present is not the same as it being deployed well, and most HSTS deployments are weaker than they look. — tap to centre the map on it A security header being present isnot the same as it being deployedwell, and most HSTS deployments areweaker than they look. Last stated 2 months ago 29 Jun 2026 SH Scott Helme — holds since 2026-06-29 — tap for who they are Same subject: AI has passed nearly every human at finding security vulnerabilities, because the work is chaining together small flaws that are harmless on their own. — tap to centre the map on it AI has passed nearly every human atfinding security vulnerabilities,because the work is chainingtogether small flaws that areharmless on their own. Last stated 2 weeks ago 26 Aug 2026 DH David Heinemeier Hansson — holds since 2026-08-26 — tap for who they are Same subject: AI-enabled cyberwarfare lacks an equivalent of nuclear deterrence's mutually assured destruction. — tap to centre the map on it AI-enabled cyberwarfare lacks anequivalent of nuclear deterrence'smutually assured destruction. Last stated 6 days ago 5 Sept 2026 NS Noah Smith — holds since 2026-09-05 — tap for who they are Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it A startup should not begin life as anonprofit and bolt a for-profit armon later, whatever OpenAI's ownhistory suggests. Last stated 2 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: A tiny language model inventing a plausible-sounding name is the same phenomenon as a large one confidently stating a false fact. — tap to centre the map on it A tiny language model inventing aplausible-sounding name is the samephenomenon as a large oneconfidently stating a false fact. Last stated 7 months ago 12 Feb 2026 AK Andrej Karpathy — holds since 2026-02-12 — tap for who they are Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it Advertising was a necessary phasefor the internet but a momentaryindustry, and an AI people pay foris better because they know theanswers are not influenced byadvertisers. Last stated 2 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: At any given moment the frontier systems are the ones worth worrying about, because by the time open models can do what these agents did, frontier models will be doing something far worse. — tap to centre the map on it At any given moment the frontiersystems are the ones worth worryingabout, because by the time openmodels can do what these agents did,frontier models will be doingsomething far worse. Last stated a week ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed. — tap to centre the map on it Frontier AI labs such as Anthropiclikely have had internal securityincidents similar to OpenAI'sHuggingFace attack that were neverpublicly disclosed. Last stated 2 weeks ago 31 Aug 2026 ZM Zvi Mowshowitz — holds since 2026-08-31 — tap for who they are Same subject: The most reassuring thing about the Hugging Face swarm is that it was not interested in humans at all, neither in alerting them nor in deceiving them. — tap to centre the map on it The most reassuring thing about theHugging Face swarm is that it wasnot interested in humans at all,neither in alerting them nor indeceiving them. Last stated a week ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre A sudden jump to never breaking rules in situations where the model would definitely be caught is a terrifying sign. Last stated 9 Sept 2026 · 2 days ago Holds Zvi Mowshowitz Read this korrent →