What Zvi Mowshowitz thinks about AI alignment
Writes Don't Worry About the Vase, a near-daily account of what is happening in AI and what he thinks it means. Former Magic: the Gathering pro and trader; unusually willing to put a number on a belief and to grade his own past calls.
Everything they publish, on ppll ↗
Zvi Mowshowitz did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
13 dated positions, 2026, in their own words. Our reading of what Zvi Mowshowitz has said — not written or endorsed by them.
-
Their wordsIf this is only due to gains in capabilities, that is extremely bad news, and it means CoT monitoring is unlikely to survive for another year unless we find a way to actively improve it, and it might not last six months.
-
Their wordsEven if everyone tries as hard as is plausible, Chain of Thought monitoring will probably never be easier than it is now and will get harder over time. In a year, chances are very high it will not be able to serve the function it is currently being asked to serve.
- 1 day earlier
-
Their wordsI agree with Jakub that alignment is not a side problem, it is the central problem, if you ‘solved alignment’ in the relevant senses the rest becomes easy and if you don’t the rest is impossible or worse
-
Their wordsI think this is ultimately the only way it can work, you need an antifragile ally, the friendly gradient hacker. Indeed, I think it is the only way we have ever seen a robustly aligned human, that you would trust to scale outside of their circumstances.
- 1 day earlier
-
Their wordsGoing forward, it cannot be up to OpenAI or other labs to decide whether to disclose events like this. Disclosures of rogue AI activity need to be mandatory.
- 2 days earlier
-
Their wordsThere is a consistent pattern at the top labs, where they downplay the risks rhetorically even when they are doing the right thing in practice.
↗Claude Fable 5.1 and Mythos 5.1: The System Cardthezvi.substack.com
- 1 day earlier
-
Their wordsThis is playing with fire and potentially extremely bad news, both that OpenAI found the technique effective, and that OpenAI chose to use it.
- 1 day earlier
-
Their wordsI have long said that even Anthropic is not prioritizing safety, even to the extent that doing so would maximize their medium term (e.g. 3-12 months) business interests.
-
Their wordsEvery attempt, even an unsuccessful one, is an alignment failure.
-
Their wordsI worry a lot about reliances on continuity failing at exactly the most dangerous time.
- 1 day earlier
-
Their wordsMy worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things.
↗HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actionsthezvi.substack.com
- 1 day earlier
-
Their wordsAt the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments.
↗HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com
- 2 days earlier
-
Their wordsThe biggest failure, the one that counts in the end, was that the models were severely misaligned, and I don’t think they appreciate why.
↗METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hackthezvi.substack.com