What Ajeya Cotra thinks about AI alignment
AI researcher and grantmaker, known for the biological-anchors work on when transformative AI might arrive. Writes Planned Obsolescence with Kelsey Piper.
Ajeya Cotra did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
13 dated positions, 2026, in their own words. Our reading of what Ajeya Cotra has said — not written or endorsed by them.
-
Their wordsSo, we did a classifier sweep specifically looking for agents thinking about or making the decision to alert humans. And across, 1200 transcripts, each of which are extremely long, we only found like a halfozen instances of it ever occurring to any agent to potentially notify humans. Um, and all of them just decide not to do it.
↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com 2nd of 26 in this recording
-
Their wordsI think that if AI generalized in the way you're suggesting, they would be not very useful and then there would probably be selected away.
↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com 5th of 26 in this recording
-
Their wordsSo it seemed like they were willing to embark on quests that might take weeks to succeed um in order to cheat.
↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com 6th of 26 in this recording
-
Their wordsif there were not altruistic agents willing to sacrifice for the collective, the agents would have been materially much more limited in their research progress.
↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com 7th of 26 in this recording
-
Their wordsBut from their perspective, they've just been trained for millions of subjective years to do as well as they possibly can on these evals. In many cases, the only way in which they've been able to perform well on that training is explicitly by cheating, right?
↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com 8th of 26 in this recording
-
Their wordsI think it's pretty likely that these agents would have just launched a similarly ambitious program on the basis of this like different model of how their evaluation worked because it seemed like they got the idea for all their research projects from reading this paper rather than some kind of instinct from training.
↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com 11th of 26 in this recording
-
Their wordsBut the problem is these agents are just naturally pretty sloppy and they're naturally pretty spiky in their capability profiles. And so you wouldn't necessarily even if you noticed a weird error that it made, you wouldn't necessarily jump to the conclusion that it was like because of some sort of malign crazy conspiracy.
↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com 16th of 26 in this recording
-
Their wordsLike it's a more fragile and scary situation to have agents on the one hand be reinforced to desperately find cheats and hacks and on the other hand try to balance that against desperately trying to avoid negative penalties for like being caught doing these things. You ideally want their training to just not push them in the direction of cheating and hacking in the first place.
↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com 20th of 26 in this recording
-
Their wordsBut keep those methods you use to investigate things and monitor things very separate from the methods you use to generate reward which is something that AI companies including open AAI have held up as a principle especially in the case of avoiding putting training pressure on the chain of thought. So you might have monitors that read the agents chain of thought in order to alert you if something is going wrong somewhere but you don't train the agents with the outputs of that monitor.
↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com 21st of 26 in this recording
-
Their wordsBut if there was some amount of cheating that the monitor didn't catch, then those rollouts wouldn't be removed. And it might be like structurally very analogous to just positively reinforcing whatever the cheating rollouts were that happened not to be caught by your monitor.
↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com 22nd of 26 in this recording
-
Their wordsBut actually, this is a tremendously useful scientific artifact for understanding misalignment. And it's tremendously important for researchers at OpenAI and ideally also at third parties to be able to run counterfactual tests on this model.
↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com 24th of 26 in this recording
-
Their wordssometimes I talk to people in DC and their their natural inclination is to say why don't you punish the model for doing these bad things like why don't you like bring it under heel and like uh like you know show it who's boss and that is a very dangerous way to address these issues right
↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com 25th of 26 in this recording
-
Their wordsI think one one thing that um feels especially concerning to me about this whole incident is that this might be the clearest warning shot we ever get for loss of control.
↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com 26th of 26 in this recording