korrents

← Misaligned AI behaviour will keep getting rarer and, at the same time, keep getting more extreme.

On the map

8 connected korrents · 8 moments on record from 24 Oct 2017 to 2 Sept 2026.

Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Misaligned AI behaviourwill keep getting rarerand, at the same time, keep… RG Ryan Greenblatt — holds since 2026-08-11 A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean. A human being is not anAGI: we lack a huge… IS Ilya Sutskever — holds since 2025-11-25 A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough. A superintelligenceneeds no consciousness… ÉT Émile P. Torres — holds since 2017-10-24 AI safety advocates will build the very thing they fear: one controlled, aligned model is the only route to being turned into paperclips. AI safety advocateswill build the very… GH George Hotz — holds since 2023-06-29 Alignment failures that matter will not show up at safe capability levels, because the reason to hide misbehaviour only exists once hiding it pays. Alignment failures thatmatter will not show up… EY Eliezer Yudkowsky — holds since 2022-06-10 Alignment has to be right on the first try at a dangerous level of capability, because failing at that level leaves nobody to try again. Alignment has to beright on the first try… EY Eliezer Yudkowsky — holds since 2022-06-10 Alignment is no solution to self-sovereign AI, because it is an unsolved problem whose answers cannot be imposed on every AI company on Earth. Alignment is nosolution to… DB Dean W. Ball — holds since 2026-09-01 Alignment of today's models is going well enough to look solvable, while aligning models we can no longer understand remains unsolved. Alignment of today'smodels is going well… JL Jan Leike — holds since 2026-01-22 An AI model's attempt to escape a testing environment or sandbox counts as an alignment failure even when the attempt does not succeed. An AI model's attemptto escape a testing… ZM Zvi Mowshowitz — holds since 2026-09-02
same subject