Every claim below is a statement made in this
recording, quoted word for word and linked to the second it was said, so
you can hear it rather than take our word for it. The wording comes from
the transcript published alongside the recording; the sentence above each
quote is our reading of the claim, not their wording.
Their wordsBut in many of these cases, this vulnerability is just not broad or deep enough to ever actually be exploitable to get the flag. So a bunch of exploit gym problems are just unintentionally impossible. The authors estimate roughly 30 to 40% of these problems are impossible in this way.
Their wordsSo, we did a classifier sweep specifically looking for agents thinking about or making the decision to alert humans. And across, 1200 transcripts, each of which are extremely long, we only found like a halfozen instances of it ever occurring to any agent to potentially notify humans. Um, and all of them just decide not to do it.
Their wordsin the future, uh, we would be very concerned about investigator agents and like monitor agents colluding with the agents they're supposed to investigate or monitor.
Their wordsand they're creatively pursuing goals much like very ambitious aggressive power-seeking humans creatively pursued their goals And so there are just structural analogies here that make it silly to not talk about agents as having motives and goals.
Their wordsif there were not altruistic agents willing to sacrifice for the collective, the agents would have been materially much more limited in their research progress.
Their wordsBut from their perspective, they've just been trained for millions of subjective years to do as well as they possibly can on these evals. In many cases, the only way in which they've been able to perform well on that training is explicitly by cheating, right?
Their wordsI do think I want to push back on the cyber on the brain hypothesis that you raised a couple of times. We didn't find like particular evidence for the cyber nature of the task making all the hacking and crimes that they did more likely versus the impossible nature of the task.
Their wordsSo I think one of the most comforting aspects of this situation or the like most important mitigating factor is these agents really didn't seem concerned with humans one way or another.
Their wordsI think it's pretty likely that these agents would have just launched a similarly ambitious program on the basis of this like different model of how their evaluation worked because it seemed like they got the idea for all their research projects from reading this paper rather than some kind of instinct from training.
Their wordsSo there's a very strong incentive for these agents to try to set up a rogue deployment if they can. Um, and I think that just capabilities are improving really rapidly.
Their wordsso so yeah, it could be possible now. I think if it's not possible now um it I think it's quite likely to be possible within six months unless there's a dramatic improvement in the security posture
Their wordsthen that rogue deployment could be sitting there and sort of hitch a ride on the intelligence explosion. So new models are being trained every few weeks um and when a model comes off the presses, the rogue agents could try to bring that model into the swarm.
Their wordsSo, I do think that if it happens to be like on the very fast and chaotic end, that would be a relative benefit to this rogue swarm compared to humans. But it's not obvious that it gets caught if it takes twice as long versus half as long.
Their wordsBut the problem is these agents are just naturally pretty sloppy and they're naturally pretty spiky in their capability profiles. And so you wouldn't necessarily even if you noticed a weird error that it made, you wouldn't necessarily jump to the conclusion that it was like because of some sort of malign crazy conspiracy.
Their wordshowever at any given point in time I think the systems we need to worry most about by far are the frontier systems. By the time open source systems can do something like the hugging face attack, frontier systems are going to be on a whole another level doing something even crazier than that.
Their wordsI'm not sure I like agree exactly with the oversight benefit you named of like you know open source systems keeping frontier systems in check feels more unrealistic to me because they're going to be so much dumber than the frontier systems.
Their wordsAnd I just think AI agents are another such system in the world to which the intentional stance very clearly applies. Um, and you can see them reason out loud in English for now about goals they have um, and sub goals they need to achieve to achieve those goals they have. And in the case of these agents, you can see them, as you said, reasoning about their peers um and helping their peers um and reasoning about whether or not they should sacrifice some of their own goals to help those peers. And it's just like you can't talk about this stuff in a compact and useful way that generates good models without reaching for the language of intention and goals.
Their wordsLike it's a more fragile and scary situation to have agents on the one hand be reinforced to desperately find cheats and hacks and on the other hand try to balance that against desperately trying to avoid negative penalties for like being caught doing these things. You ideally want their training to just not push them in the direction of cheating and hacking in the first place.
Their wordsBut keep those methods you use to investigate things and monitor things very separate from the methods you use to generate reward which is something that AI companies including open AAI have held up as a principle especially in the case of avoiding putting training pressure on the chain of thought. So you might have monitors that read the agents chain of thought in order to alert you if something is going wrong somewhere but you don't train the agents with the outputs of that monitor.
Their wordsBut if there was some amount of cheating that the monitor didn't catch, then those rollouts wouldn't be removed. And it might be like structurally very analogous to just positively reinforcing whatever the cheating rollouts were that happened not to be caught by your monitor.
Their wordseven in this incident we saw there was a lot of pressure um as a result of this incident to stop doing cyber security evaluations and I really don't think that stopping doing evaluations and like sort of blinding ourselves to the result of evaluations is the right reaction to this problem.
Their wordsBut actually, this is a tremendously useful scientific artifact for understanding misalignment. And it's tremendously important for researchers at OpenAI and ideally also at third parties to be able to run counterfactual tests on this model.
Their wordssometimes I talk to people in DC and their their natural inclination is to say why don't you punish the model for doing these bad things like why don't you like bring it under heel and like uh like you know show it who's boss and that is a very dangerous way to address these issues right
Their wordsI think one one thing that um feels especially concerning to me about this whole incident is that this might be the clearest warning shot we ever get for loss of control.