korrents

korrents · Dwarkesh Podcast

Dylan Patel — The single biggest bottleneck to scaling AI compute

Dylan Patel · 2h 30m · youtube.com

28 korrents from this recording

1h2h
Dylan Patel did not write this page.

Every claim below is a statement made in this recording, quoted word for word and linked to the second it was said, so you can hear it rather than take our word for it. The wording comes from the transcript published alongside the recording; the sentence above each quote is our reading of the claim, not their wording.

  1. 0:15:28 · watch on youtube.com

    if improvement stopped you know here the value of an H100 is now predicated on the value that GPD 5.4 four can get out of it instead of the value that GP4 can get out of it and the margins and all that stuff that these labs are doing and they're in a competitive environment so their margins can't go to infinity. Um so you sort of have this like dynamic that is quite interesting in that an H100 is worth more today than it was 3 years ago.
  2. 6 min later
  3. 0:21:32 · watch on youtube.com

    A companies that have locked up, you know, and and don't have commitment issues, you know, have these 5-year contracts for compute, they've kind of locked in a humongous margin advantage because they've locked in compute for 5 years at a price of what it transacted at 5 years ago or three years ago or two years ago, whatever it is.
  4. 3 min later
  5. 0:24:36 · watch on youtube.com

    Um I think at least this year we're going to see margins for the model vendors go up a lot, right? Because they're so capacity constrained, they have to demand destroy demand, right? there is there's no way they can continue anthropic can continue at the current pace without destroying demand.
  6. 3 min later
  7. 0:27:20 · watch on youtube.com

    TSMC is much more excited to give allocation to Graviton than they are to tranium because they view CPU business as more stable long-term growth right and as a company that is conservative and doesn't want to ride cycles of growth too hard you actually want to allocate to the uh the market that is more stable and lower growth rate first before you allocate all the incremental capacity to the fast growth rate market.
  8. 5 min later
  9. 0:32:25 · watch on youtube.com

    Because Enthropic saw it before Google. And then Google had Nano Banano and Gemini 3 which caused their user metrics to skyrocket and leadership at Google was like oh and then they started making the statement of we have to double compute every is it 6 months or I don't remember the exact number that they said.
  10. 2 min later
  11. 0:34:51 · watch on youtube.com

    Yeah, I think the biggest bottleneck is compute and for that the longest lead time supply chains are not power or data centers. They're actually the semiconductor supply chain themselves, right? It switches back from being power and data center uh as a major bottleneck to chips.
  12. 2 min later
  13. 0:37:03 · watch on youtube.com

    So to scale compute further right there's some different bottlenecks this year next year uh but ultimately by 2829 the bottleneck falls to the lowest rung on the supply chain which is ASML right ASML makes the world's most complicated machine i.e. an EUV tool.
  14. 4 min later
  15. 0:40:48 · watch on youtube.com

    oh 50 gigawatts of economic you know sort of capex in in the data center and what gets built on top of that in terms of tokens is even larger right it might be hundred billion dollars worth of AI value into the supply chain is held up by this $1.2 two billion dollars worth of tooling that simply just cannot expand its supply chain quickly.
  16. 2 min later
  17. 0:42:25 · watch on youtube.com

    Um, and then you stack on 70 this year, 80 next year, growing to 100 by 2030. You're at like 700 EV tools by the end of the decade. Um, 700 EV tools, three and a half tools per gigawatt. um assuming it's all allocated to AI which it's not but three and a half tools per gigawatt gets you to 200 gigawatts worth of AI chips for the data centers to deploy
  18. 3 min later
  19. 0:45:42 · watch on youtube.com

    You can you can take the margin like Nvidia takes the margin. memory players are taking the margin, but ASML has never risen the price more than they've increased the capability of the tool. Um, and so in a sense, they've always provided net benefit to their customer.
  20. 1 min later
  21. 0:46:35 · watch on youtube.com

    Um in general the semiconductor supply chain has not right it's lived through the booms and bust and uh we can talk a bit more about it but basically no one you know some players as of very recently have like woken up but in general no one really sees demand for 200 gawatts a year of AI chips or you know trillions of dollars of spend a year in the semiconductor supply chain.
  22. 16 min later
  23. 1:02:11 · watch on youtube.com

    So when you look at inference at let's say 100 tokens a second for deepseek and kimk 2.5 hopper versus blackwell the performance difference is on the order of 20x
  24. 5 min later
  25. 1:07:30 · watch on youtube.com

    I think they'll have working tools. I don't think that they'll be able to manufacture a bunch yet, right? You know, there's they're sort of having it work and then there's production hell, right?
  26. 4 min later
  27. 1:11:44 · watch on youtube.com

    As we move from, you know, hey, [clears throat] these companies are selling tokens where they provide the entire uh reasoning chain and all that to uh selling automated, you know, white collar work, right? Automated software engineer, send them the request, they give you the result back and there's a bunch of thinking on the back end that they don't show you. The ability to distill out of American models into Chinese models will be harder.
  28. 4 min later
  29. 1:15:51 · watch on youtube.com

    But I don't know like I don't know what fast timelines means, right? Like I I like don't think you have to believe in AGI to have the timelines where the US wins.
  30. 2 min later
  31. 1:17:38 · watch on youtube.com

    They could release claw slow mode and have an increase in tokens per dollar by a significant amount. Um they could probably like reduce the price of Opus 46 by you know 4x 5x and reduce the speed by another by maybe just like 2x like the curve on inference throughput versus speed is there already just on hm um and yet they don't um because no one actually wants to use a slow model
  32. 5 min later
  33. 1:22:09 · watch on youtube.com

    and even if you take a generous interpretation of 128 * 8 gig transfers, you're at 128 gigabytes a second for the same shoreline versus 2 and a half terabytes a second. There's a there's an order of magnitude difference in bandwidth per edge area.
  34. 5 min later
  35. 1:26:55 · watch on youtube.com

    Um DRAM gets released goes to AI chips who are willing to do longer term contracts, willing to pay higher margins, etc., etc. because at the end of the day, the margin that they extract is much larger from the end user or whatever. Um, and so this this this probably leads to like people hating AI even more, right?
  36. 4 min later
  37. 1:30:38 · watch on youtube.com

    Uh, Micron bought a fab from a company in Taiwan that makes lagging edge chips, right? Um, Heinix and Samsung are doing, you know, some pretty crazy things to try and expand capacity at their existing fabs, uh, that also have like very large knock-on effects in the economy. And so, hey, why can't we build more capacity is like there's nowhere to put the tools, right?
  38. 5 min later
  39. 1:35:12 · watch on youtube.com

    but I think he can build the uh clean room. It'll take a year or two. Maybe initially it won't be super fast, but then over time you'll get faster and faster at it. But then the really complex part is actually developing the process technology and building wafers. And I don't think he can develop that uh quickly. I think that has a lot of built-up knowledge.
  40. 12 min later
  41. 1:46:51 · watch on youtube.com

    Then all of a sudden, you've unlocked 20% of the US grid for data centers because most of the times that capacity is sitting idle and it's really only there for that peak, right? Which is a day or two, right?
  42. 5 min later
  43. 1:52:02 · watch on youtube.com

    um humanoid robots maybe start to or robotics at least start to but the main factor is going to be for reducing the number of people is modularizing things and making them in factories in Asia
  44. 5 min later
  45. 1:57:13 · watch on youtube.com

    And so I think people are figuring out how to build these things and permitting like I I just like ultimately like permitting and red tape in middle of nowhere Texas or middle of nowhere Wyoming or middle of nowhere like New Mexico is probably a hell of a lot easier than sending stuff into space
  46. 5 min later
  47. 2:02:13 · watch on youtube.com

    So, space data centers effectively are not li limited by, you know, hey, we have this energy advantage. It's actually just limited by the same contended resource. We can only make 200 gawatts of chips a year by the end of the decade.
  48. 9 min later
  49. 2:10:54 · watch on youtube.com

    well the model the compute efficiency gains you get from research are so large you actually want most of your compute to go to research not to development because you know all these researchers are generating new ideas trying them out testing them and continuing to march along this and push the prao optimal curve of scaling laws further and further and further
  50. 9 min later
  51. 2:19:46 · watch on youtube.com

    and so I don't think TSMC would kick out Apple. I think Apple will become a smaller and smaller and smaller percentage of TSMC's revenue and therefore be less relevant for TSMC to cater to their demands.
  52. 4 min later
  53. 2:23:30 · watch on youtube.com

    And Huawei has a bigger pool in China. It's very arguable that Huawei, if they had TSMC, would be better than Nvidia.
  54. 6 min later
  55. 2:29:35 · watch on youtube.com

    Um, just shipping out all the engineers and blowing up the fabs means China has a stronger semiconductor supply chain than the rest of the world, right?