What Scott Alexander thinks about AI alignment
Essayist; writes Astral Codex Ten, previously Slate Star Codex. Keeps a numbered public list of the things he has been convinced he got wrong, which is where most of the turns recorded here come from.
Everything they publish, on ppll ↗
Scott Alexander did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
5 dated positions, 2023 to 2026, in their own words. Our reading of what Scott Alexander has said — not written or endorsed by them.
-
Their wordsWe’re positively reinforcing AI for success on benchmarks, including impossible benchmarks , then negatively reinforcing it for getting caught cheating.
-
Their wordsRather than treat AI as similar to airplanes, we should treat it as something between airplanes and humans. We don’t know exactly where on this spectrum they will land, but we can no longer be certain that AI will lack any given humanlike motivation.
- 3 months earlier
-
Their wordsI think there's a 25% chance of AGI by 2027, a 50% chance by 2034, and a 75% chance by 2045.
-
Their wordsI think there's a 50% chance we get a _warning shot_ before AI crosses the point of no return.
- 3 years earlier
-
Their wordsAnd I notice that the tiny handful of people capable of caring about 200,000 people dying of neglected tropical diseases are the same tiny handful of people capable of caring about the next pandemic, or superintelligence, or human extinction.
↗In Continued Defense Of Effective Altruismastralcodexten.com 4th of 5 in this piece