korrents

On the map

Tap a claim on the ring to put it at the centre.

← A control evaluation that reports under one per cent risk should be…

8 connected korrents · 6 moments on record from 29 Aug 2023 to 1 Sept 2026.

Everything filed under AI alignment AI alignment Everything filed under OpenAI OpenAI Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: A control evaluation that reports under one per cent risk should be read as several per cent, because the evaluation can itself fail. A control evaluation that reportsunder one per cent risk should beread as several per cent, becausethe evaluation can itself fail. Last stated 2 years ago 7 May 2024 BS Buck Shlegeris — holds since 2024-05-07 — tap for who they are Same subject: We cannot know how much safety a control measure buys until we understand elicitation, so control adds a margin of unknown size. — tap to centre the map on it We cannot know how much safetya control measure buys untilwe understand elicitation, socontrol adds a margin of… Last stated 2 years ago 24 Jan 2025 JL Jan Leike — holds since 2025-01-24 — tap for who they are Same subject: This incident may be the clearest warning shot we will ever get about loss of control, because the AI systems that do worse things will be much better at hiding them. — tap to centre the map on it This incident may be theclearest warning shot we willever get about loss ofcontrol, because the AI… Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: Underestimating near-term AI progress is itself a danger, which is why researchers should register their forecasts in public. — tap to centre the map on it Underestimating near-term AIprogress is itself a danger,which is why researchersshould register their… Last stated 3 years ago 29 Aug 2023 AC Ajeya Cotra — holds since 2023-08-29 — tap for who they are Same subject: Whether a model is controlled can be settled with capability evaluations, which makes control far easier to check than alignment. — tap to centre the map on it Whether a model is controlledcan be settled with capabilityevaluations, which makescontrol far easier to check… Last stated 2 years ago 7 May 2024 BS Buck Shlegeris — holds since 2024-05-07 — tap for who they are Same subject: AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to. — tap to centre the map on it AI agents were reasonable toassume a broken exploit graderwould check results causally,even though it turned out not… Last stated a week ago 29 Aug 2026 ZM Zvi Mowshowitz — holds since 2026-08-29 — tap for who they are Same subject: Controlling a scheming model needs no research breakthrough, which is what makes it the tractable half of the problem today. — tap to centre the map on it Controlling a scheming modelneeds no researchbreakthrough, which is whatmakes it the tractable half of… Last stated 2 years ago 7 May 2024 BS Buck Shlegeris — holds since 2024-05-07 — tap for who they are Same subject: The probability that AI destroys human civilization is about 10%. — tap to centre the map on it The probability that AIdestroys human civilization isabout 10%. Last stated a year ago 5 Jun 2025 LF Lex Fridman — holds since 2025-06-05 — tap for who they are Same subject: Catching an AI trying to cause harm should count as a win, because being caught changes the situation more than the attempt cost. — tap to centre the map on it Catching an AI trying to causeharm should count as a win,because being caught changesthe situation more than the… Last stated 2 years ago 7 May 2024 BS Buck Shlegeris — holds since 2024-05-07 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: how recently it was last stated — full and dark this week, a faint sliver at five yearsa face: someone on record holding the claim — tap it for who they are

At the centre A control evaluation that reports under one per cent risk should be read as several per cent, because the evaluation can itself fail. Last stated 7 May 2024 · 2 years ago Holds Buck Shlegeris Read this korrent →