korrents

korrents · No Priors

Skill Issue: Andrej Karpathy on Code Agents, AutoResearch, and the Loopy Era of AI

Andrej Karpathy · 1h 06m · youtube.com

20 korrents from this recording

1h
Andrej Karpathy did not write this page.

Every claim below is a statement made in this recording, quoted word for word and linked to the second it was said, so you can hear it rather than take our word for it. The wording comes from the transcript published alongside the recording; the sentence above each quote is our reading of the claim, not their wording.

  1. 0:01:45 · watch on youtube.com

    I kind of went from 80/20 of like, you know, uh to like 20/80 of writing code by myself versus just delegating to agents. And I don't even think it's 20/80 by now. I think it's a lot more than that. I don't think I've typed like a line of code probably since December basically.
  2. 2 min later
  3. 0:03:36 · watch on youtube.com

    I think to a large extent you feel like it's a skill issue. It's not that the capability is not there. It's that you just haven't found a way to string it together of what's available.
  4. 2 min later
  5. 0:05:54 · watch on youtube.com

    But now it's not about flops, it's about tokens. So what is your token throughput and what token throughput do you command?
  6. 7 min later
  7. 0:13:11 · watch on youtube.com

    these apps that are on the app store for using these smart home devices, etc. Uh, these shouldn't even exist kind of in a certain sense. Like shouldn't it just be APIs and shouldn't agents be just using it directly?
  8. 1 min later
  9. 0:14:12 · watch on youtube.com

    So I think the industry just has to reconfigure in so many ways that's like the customer is not the human anymore. It's like agents who are acting on behalf of humans and this refactoring will be will probably be substantial in a certain sense.
  10. 3 min later
  11. 0:17:20 · watch on youtube.com

    So the question is how do I refactor all the abstractions so that I'm not I have to arrange it once and hit go. The name of the game is how can you get more agents running for longer periods of time without your involvement doing stuff on your behalf?
  12. 1 min later
  13. 0:18:38 · watch on youtube.com

    And I've gotten to a certain point and I thought it was like fairly well tuned and then I let auto research go for like overnight and it came back with like tunings that I didn't see.
  14. 2 min later
  15. 0:20:11 · watch on youtube.com

    But okay, they shouldn't actually be enacting these ideas. There is a queue of ideas and there's maybe an automated scientist that comes up with ideas based on all the archive papers and GitHub repos and it funnels ideas in or researchers can contribute ideas, but it's a single queue and there is workers that pull items and they try them out.
  16. 4 min later
  17. 0:24:37 · watch on youtube.com

    I simultaneously feel like I'm talking to an extremely brilliant PhD student who's been like a systems programmer for their entire life and a 10-year-old.
  18. 2 min later
  19. 0:26:07 · watch on youtube.com

    And so you're kind of like you're either on rails and you're part of the super intelligence circuits or you're not on rails and you're outside of the verifiable domains and suddenly everything kind of just like meanders.
  20. 1 min later
  21. 0:26:52 · watch on youtube.com

    So even though the models have improved tremendously and if you give them an agentic task, they will just go for hours and move mountains for you. And then you ask for like a joke and it has a stupid joke. It's crappy joke from five years ago and it's because it's outside of the it's outside of the RL.
  22. 3 min later
  23. 0:29:42 · watch on youtube.com

    I do think we should expect more speciation in the intelligences. Um like, you know, the animal kingdom is extremely diverse in the brains that exist and there's lots of different niches of of nature and some animals have overdeveloped visual cortex or other part kind of parts and I think we we should be able to see more speciation and um you don't need like this oracle that knows everything.
  24. 6 min later
  25. 0:36:01 · watch on youtube.com

    a swarm of agents on the internet could collaborate to improve LLMs and could potentially even like run circles around frontier labs. Like who knows, you know? Um yeah, like maybe that's even possible. Like frontier labs have a huge amount of trusted compute but the earth is much bigger and has huge amount of untrusted compute.
  26. 7 min later
  27. 0:42:57 · watch on youtube.com

    I do have like cautiously optimistic view of this in software engineering where I do think um it does seem to me like the demand for software will be extremely large. Um and it's just become a lot cheaper.
  28. 3 min later
  29. 0:45:51 · watch on youtube.com

    You're you're not a completely free agent and you can't actually like be part of that conversation in a fully autonomous um free way. Like if you're inside one of the frontier labs. Like there's some things that you can't say. Uh and conversely there are some things that the organization wants you to say.
  30. 2 min later
  31. 0:47:24 · watch on youtube.com

    And I think if you're outside of that frontier lab, your your judgment fundamentally will start to drift because you're not part of the you know, what's coming down the line. And so I feel like my judgment will inevitably start to drift as well.
  32. 5 min later
  33. 0:52:20 · watch on youtube.com

    But I want there to be a thing that's behind and that uh is kind of like a common working space for intelligences that the entire industry has access to.
  34. 4 min later
  35. 0:55:50 · watch on youtube.com

    you're going to run out of things that you're going to do purely in the digital space. At some point you have to go to the universe and you have to ask it questions. Um you have to run an experiment and see what the universe tells you to get back to learn something.
  36. 9 min later
  37. 1:04:40 · watch on youtube.com

    It used to be that you have documentation for other people who are going to use your library, but like you shouldn't do that anymore. Like you should have instead of HTML documents for humans, you have markdown documents for agents.
  38. 1:05:05 · watch on youtube.com

    so for example, micro GPT, like I asked I tried to get an agent to write micro GPT. So, I told it like try to boil down the simplest things. Like try to boil down my um neural network training to the simplest thing and it can't do it.