korrents

korrents · Dwarkesh Podcast

Ajeya Cotra – "This might be the clearest warning shot we ever get"

Ajeya Cotra · 2h 20m · youtube.com

26 korrents from this recording

1h2h
Ajeya Cotra did not write this page.

Every claim below is a statement made in this recording, quoted word for word and linked to the second it was said, so you can hear it rather than take our word for it. The wording comes from the transcript published alongside the recording; the sentence above each quote is our reading of the claim, not their wording.

  1. 0:00:46 · watch on youtube.com

    But in many of these cases, this vulnerability is just not broad or deep enough to ever actually be exploitable to get the flag. So a bunch of exploit gym problems are just unintentionally impossible. The authors estimate roughly 30 to 40% of these problems are impossible in this way.
  2. 32 min later
  3. 0:33:02 · watch on youtube.com

    So, we did a classifier sweep specifically looking for agents thinking about or making the decision to alert humans. And across, 1200 transcripts, each of which are extremely long, we only found like a halfozen instances of it ever occurring to any agent to potentially notify humans. Um, and all of them just decide not to do it.
  4. 17 min later
  5. 0:50:23 · watch on youtube.com

    in the future, uh, we would be very concerned about investigator agents and like monitor agents colluding with the agents they're supposed to investigate or monitor.
  6. 4 min later
  7. 0:54:15 · watch on youtube.com

    and they're creatively pursuing goals much like very ambitious aggressive power-seeking humans creatively pursued their goals And so there are just structural analogies here that make it silly to not talk about agents as having motives and goals.
  8. 6 min later
  9. 0:59:53 · watch on youtube.com

    I think that if AI generalized in the way you're suggesting, they would be not very useful and then there would probably be selected away.
  10. 3 min later
  11. 1:02:59 · watch on youtube.com

    So it seemed like they were willing to embark on quests that might take weeks to succeed um in order to cheat.
  12. 1 min later
  13. 1:04:16 · watch on youtube.com

    if there were not altruistic agents willing to sacrifice for the collective, the agents would have been materially much more limited in their research progress.
  14. 3 min later
  15. 1:07:16 · watch on youtube.com

    But from their perspective, they've just been trained for millions of subjective years to do as well as they possibly can on these evals. In many cases, the only way in which they've been able to perform well on that training is explicitly by cheating, right?
  16. 3 min later
  17. 1:09:54 · watch on youtube.com

    I do think I want to push back on the cyber on the brain hypothesis that you raised a couple of times. We didn't find like particular evidence for the cyber nature of the task making all the hacking and crimes that they did more likely versus the impossible nature of the task.
  18. 5 min later
  19. 1:14:56 · watch on youtube.com

    So I think one of the most comforting aspects of this situation or the like most important mitigating factor is these agents really didn't seem concerned with humans one way or another.
  20. 2 min later
  21. 1:17:22 · watch on youtube.com

    I think it's pretty likely that these agents would have just launched a similarly ambitious program on the basis of this like different model of how their evaluation worked because it seemed like they got the idea for all their research projects from reading this paper rather than some kind of instinct from training.
  22. 8 min later
  23. 1:24:53 · watch on youtube.com

    So there's a very strong incentive for these agents to try to set up a rogue deployment if they can. Um, and I think that just capabilities are improving really rapidly.
  24. 3 min later
  25. 1:27:44 · watch on youtube.com

    so so yeah, it could be possible now. I think if it's not possible now um it I think it's quite likely to be possible within six months unless there's a dramatic improvement in the security posture
  26. 2 min later
  27. 1:29:32 · watch on youtube.com

    then that rogue deployment could be sitting there and sort of hitch a ride on the intelligence explosion. So new models are being trained every few weeks um and when a model comes off the presses, the rogue agents could try to bring that model into the swarm.
  28. 4 min later
  29. 1:33:36 · watch on youtube.com

    So, I do think that if it happens to be like on the very fast and chaotic end, that would be a relative benefit to this rogue swarm compared to humans. But it's not obvious that it gets caught if it takes twice as long versus half as long.
  30. 3 min later
  31. 1:36:33 · watch on youtube.com

    But the problem is these agents are just naturally pretty sloppy and they're naturally pretty spiky in their capability profiles. And so you wouldn't necessarily even if you noticed a weird error that it made, you wouldn't necessarily jump to the conclusion that it was like because of some sort of malign crazy conspiracy.
  32. 3 min later
  33. 1:39:53 · watch on youtube.com

    however at any given point in time I think the systems we need to worry most about by far are the frontier systems. By the time open source systems can do something like the hugging face attack, frontier systems are going to be on a whole another level doing something even crazier than that.
  34. 1 min later
  35. 1:41:13 · watch on youtube.com

    I'm not sure I like agree exactly with the oversight benefit you named of like you know open source systems keeping frontier systems in check feels more unrealistic to me because they're going to be so much dumber than the frontier systems.
  36. 9 min later
  37. 1:50:34 · watch on youtube.com

    And I just think AI agents are another such system in the world to which the intentional stance very clearly applies. Um, and you can see them reason out loud in English for now about goals they have um, and sub goals they need to achieve to achieve those goals they have. And in the case of these agents, you can see them, as you said, reasoning about their peers um and helping their peers um and reasoning about whether or not they should sacrifice some of their own goals to help those peers. And it's just like you can't talk about this stuff in a compact and useful way that generates good models without reaching for the language of intention and goals.
  38. 6 min later
  39. 1:56:09 · watch on youtube.com

    Like it's a more fragile and scary situation to have agents on the one hand be reinforced to desperately find cheats and hacks and on the other hand try to balance that against desperately trying to avoid negative penalties for like being caught doing these things. You ideally want their training to just not push them in the direction of cheating and hacking in the first place.
  40. 1 min later
  41. 1:57:02 · watch on youtube.com

    But keep those methods you use to investigate things and monitor things very separate from the methods you use to generate reward which is something that AI companies including open AAI have held up as a principle especially in the case of avoiding putting training pressure on the chain of thought. So you might have monitors that read the agents chain of thought in order to alert you if something is going wrong somewhere but you don't train the agents with the outputs of that monitor.
  42. 2 min later
  43. 1:59:23 · watch on youtube.com

    But if there was some amount of cheating that the monitor didn't catch, then those rollouts wouldn't be removed. And it might be like structurally very analogous to just positively reinforcing whatever the cheating rollouts were that happened not to be caught by your monitor.
  44. 8 min later
  45. 2:07:18 · watch on youtube.com

    even in this incident we saw there was a lot of pressure um as a result of this incident to stop doing cyber security evaluations and I really don't think that stopping doing evaluations and like sort of blinding ourselves to the result of evaluations is the right reaction to this problem.
  46. 1 min later
  47. 2:08:11 · watch on youtube.com

    But actually, this is a tremendously useful scientific artifact for understanding misalignment. And it's tremendously important for researchers at OpenAI and ideally also at third parties to be able to run counterfactual tests on this model.
  48. 1 min later
  49. 2:08:44 · watch on youtube.com

    sometimes I talk to people in DC and their their natural inclination is to say why don't you punish the model for doing these bad things like why don't you like bring it under heel and like uh like you know show it who's boss and that is a very dangerous way to address these issues right
  50. 7 min later
  51. 2:16:04 · watch on youtube.com

    I think one one thing that um feels especially concerning to me about this whole incident is that this might be the clearest warning shot we ever get for loss of control.