korrents

On the map

Tap a claim on the ring to put it at the centre.

← Alignment failures that matter will not show up at safe capability…

8 connected korrents · 8 moments on record from 24 Oct 2017 to 2 Sept 2026.

Everything filed under AI alignment AI alignment Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Alignment failures that matter will not show up at safe capability levels, because the reason to hide misbehaviour only exists once hiding it pays. Alignment failures that matter willnot show up at safe capabilitylevels, because the reason to hidemisbehaviour only exists once hidingit pays. Last stated 4 years ago 10 Jun 2022 EY Eliezer Yudkowsky — holds since 2022-06-10 — tap for who they are Same subject: A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean. — tap to centre the map on it A human being is not an AGI:we lack a huge amount ofknowledge and rely oncontinual learning instead, so… Last stated 9 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough. — tap to centre the map on it A superintelligence needs noconsciousness, emotions ormalice to be dangerous; a goalsystem slightly misaligned… Last stated 9 years ago 24 Oct 2017 ÉT Émile P. Torres — holds since 2017-10-24 — tap for who they are Same subject: AI safety advocates will build the very thing they fear: one controlled, aligned model is the only route to being turned into paperclips. — tap to centre the map on it AI safety advocates will buildthe very thing they fear: onecontrolled, aligned model isthe only route to being turned… Last stated 3 years ago 29 Jun 2023 GH George Hotz — holds since 2023-06-29 — tap for who they are Same subject: Alignment has to be right on the first try at a dangerous level of capability, because failing at that level leaves nobody to try again. — tap to centre the map on it Alignment has to be right onthe first try at a dangerouslevel of capability, becausefailing at that level leaves… Last stated 4 years ago 10 Jun 2022 EY Eliezer Yudkowsky — holds since 2022-06-10 — tap for who they are Same subject: Alignment is no solution to self-sovereign AI, because it is an unsolved problem whose answers cannot be imposed on every AI company on Earth. — tap to centre the map on it Alignment is no solution toself-sovereign AI, because itis an unsolved problem whoseanswers cannot be imposed on… Last stated 6 days ago 1 Sept 2026 DB Dean W. Ball — holds since 2026-09-01 — tap for who they are Same subject: Alignment of today's models is going well enough to look solvable, while aligning models we can no longer understand remains unsolved. — tap to centre the map on it Alignment of today's models isgoing well enough to looksolvable, while aligningmodels we can no longer… Last stated 7 months ago 22 Jan 2026 JL Jan Leike — holds since 2026-01-22 — tap for who they are Same subject: An AI model's attempt to escape a testing environment or sandbox counts as an alignment failure even when the attempt does not succeed. — tap to centre the map on it An AI model's attempt toescape a testing environmentor sandbox counts as analignment failure even when… Last stated 5 days ago 2 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-02 — tap for who they are Same subject: Arguing about whether superintelligence arrives in two years or five is close to a waste of time, because very powerful models are inevitable either way. — tap to centre the map on it Arguing about whethersuperintelligence arrives intwo years or five is close toa waste of time, because very… Last stated a month ago 29 Jul 2026 AW Alexandr Wang — holds since 2026-07-29 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: how recently it was last stated — full and dark this week, a faint sliver at five yearsa face: someone on record holding the claim — tap it for who they are

At the centre Alignment failures that matter will not show up at safe capability levels, because the reason to hide misbehaviour only exists once hiding it pays. Last stated 10 Jun 2022 · 4 years ago Holds Eliezer Yudkowsky Read this korrent →