korrents
Andrej Karpathy

What Andrej Karpathy thinks about reinforcement learning

@andrej-karpathy · 66 positions · 0 changes of mind

Founding member of OpenAI and former director of AI at Tesla; creator of nanoGPT and the term "vibe coding".

Everything they publish, on ppll ↗

Andrej Karpathy did not write this page.

We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.

11 dated positions, 2025 to 2026, in their own words. Our reading of what Andrej Karpathy has said — not written or endorsed by them.

  1. And I've gotten to a certain point and I thought it was like fairly well tuned and then I let auto research go for like overnight and it came back with like tunings that I didn't see.

    Skill Issue: Andrej Karpathy on Code Agents, AutoResearch, and the Loopy Era of AIyoutube.com 7th of 20 in this recording

  2. And so you're kind of like you're either on rails and you're part of the super intelligence circuits or you're not on rails and you're outside of the verifiable domains and suddenly everything kind of just like meanders.

    Skill Issue: Andrej Karpathy on Code Agents, AutoResearch, and the Loopy Era of AIyoutube.com 10th of 20 in this recording

    AI alignment

  3. So even though the models have improved tremendously and if you give them an agentic task, they will just go for hours and move mountains for you. And then you ask for like a joke and it has a stupid joke. It's crappy joke from five years ago and it's because it's outside of the it's outside of the RL.

    Skill Issue: Andrej Karpathy on Code Agents, AutoResearch, and the Loopy Era of AIyoutube.com 11th of 20 in this recording

  4. 4 months earlier
  5. If a task/job is verifiable, then it is optimizable directly or via reinforcement learning, and a neural net can be trained to work extremely well.

    Verifiabilitykarpathy.bearblog.dev

  6. 4 weeks earlier
  7. And it just so turns out that um this was extremely early, way too early. so early that we shouldn't have been working on that, you know, uh because um if you're just stumbling your way around and keyboard mashing and mouse clicking and trying to get rewards in these environments, um your reward is too sparse and you just won't learn and you're going to burn a forest uh computing and you're never actually going to get something off the ground.

    Andrej Karpathy — “We’re summoning ghosts, not building animals”youtube.com 2nd of 30 in this recording

  8. a lot of what looks like learning is actually a lot more maturation of the brain and I think that actually very little reinforcement learning for animals and I think a lot of the reinforcement learning is actually like more like motor tasks. It's not intelligence tasks. So I actually kind of think humans don't actually like really use RL roughly speaking is what I would say.

    Andrej Karpathy — “We’re summoning ghosts, not building animals”youtube.com 4th of 30 in this recording

  9. So that's why I kind of call pre-training this kind of like crappy evolution. It's like the practically possible version with our technology and what we have available to us to get to a starting point where we can actually do things like reinforcement learning and so on.

    Andrej Karpathy — “We’re summoning ghosts, not building animals”youtube.com 5th of 30 in this recording

    scaling laws

  10. reinforcement learning is a lot worse than I think the average person thinks reinforcement learning is terrible. It just so happens that uh everything that we had before is much worse

    Andrej Karpathy — “We’re summoning ghosts, not building animals”youtube.com 12th of 30 in this recording

  11. A human would never do this. Number one, a human would never do hundreds of rollouts. Uh number two, when a person sort of finds a solution, they will have a pretty complicated process of review of like, okay, I think these parts that I did well, these parts I did not do that well, I should probably do this or that.

    Andrej Karpathy — “We’re summoning ghosts, not building animals”youtube.com 13th of 30 in this recording

  12. the reason that I think this is kind of tricky is quite subtle. And it's the fact that anytime you use an LLM to assign a reward, those LLMs are giant things with billions of parameters and they're gameable.

    Andrej Karpathy — “We’re summoning ghosts, not building animals”youtube.com 14th of 30 in this recording

    LLMs

  13. It's about booting up a brain and I think physics uniquely boots up the brain the best.

    Andrej Karpathy — “We’re summoning ghosts, not building animals”youtube.com 30th of 30 in this recording

    physicseducation