korrents

korrents · Lex Fridman Podcast · #459

DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459

Nathan Lambert, Dylan Patel · 5h 06m · youtube.com

44 korrents from this recording

1h2h3h4h5h
Nathan Lambert, Dylan Patel did not write this page.

Every claim below is a statement made in this recording, quoted word for word and linked to the second it was said, so you can hear it rather than take our word for it. The wording comes from the transcript published alongside the recording; the sentence above each quote is our reading of the claim, not their wording.

  1. 0:06:17 · watch on youtube.com · Nathan Lambert

    And we'll get into the details of the models and again and again as we try to get deeper into how the models were trained, we will say things like the data processing, data filtering data quality is the number one determinant of the model quality.
  2. 34 min later
  3. 0:40:20 · watch on youtube.com · Nathan Lambert

    The scale word gets a lot of attention in this. The interpretation that I use is effectively to avoid adding the human priors to your learning process. And if you read the original essay, this is what it talks about is how researchers will try to come up with clever solutions to their specific problem that might get them small gains in the short term while simply enabling these deep learning systems to work efficiently, and for these bigger problems in the long term might be more likely to scale and continue to drive success.
  4. 9 min later
  5. 0:49:00 · watch on youtube.com · Nathan Lambert

    This is why you want to work in post-training because the GPU cost for training is lower. So you can make a higher percentage of your training runs YOLO runs.
  6. 1 min later
  7. 0:50:25 · watch on youtube.com · Dylan Patel

    I think it's even more impressive what OpenAI did in 2022. At the time, no one believed in mixture of experts models at Google who had all the researchers. OpenAI had such little compute and they devoted all of their compute for many months, all of it, 100% for many months to GPT-4 with a brand-new architecture with no belief that, "Hey, let me spend a couple of hundred million dollars, which is all of the money I have on this model." That is truly YOLO.
  8. 6 min later
  9. 0:56:33 · watch on youtube.com · Nathan Lambert

    Accepted practice is that for any given model that is a notable advancement, you're going to do two to 4x compute of the full training run in experiments alone.
  10. 5 min later
  11. 1:02:02 · watch on youtube.com · Nathan Lambert

    There's not many worlds where China cannot train AI models. I think export controls are decapping the amount of compute or the density of compute that China can have.
  12. 1 min later
  13. 1:03:30 · watch on youtube.com · Dylan Patel

    To some extent, training a model does effectively nothing. They have a model. The thing that Dario is sort of speaking to is the implementation of that model, once trained to then create huge economic growth, huge increases in military capabilities, huge increases in productivity of people, betterment of lives.
  14. 5 min later
  15. 1:08:05 · watch on youtube.com · Nathan Lambert

    I think my personal definition of AGI is much simpler. I think language models are a form of AGI and all of this super powerful stuff is a next step that's great if we get these tools. But a language model has so much value in so many domains that it's a general intelligence to me.
  16. 5 min later
  17. 1:12:49 · watch on youtube.com · Nathan Lambert

    There's some research that shows that the distribution is actually the limiting factor. So language models haven't yet made misinformation particularly change the equation there.
  18. 6 min later
  19. 1:18:56 · watch on youtube.com · Dylan Patel

    if you believe we're in this sort of stage of economic growth and change that we've been in for the last 20 years, the export controls are absolutely guaranteeing that China will win long-term.
  20. 1 min later
  21. 1:20:02 · watch on youtube.com · Dylan Patel

    China, if they wanted to build the largest data center in the world, if they had access to the chips, could. So it's just a question of when, not if.
  22. 23 min later
  23. 1:43:24 · watch on youtube.com · Dylan Patel

    Arizona is a paperweight. If Hsinchu disappeared off the face of the planet, within a year, couple years, Arizona would stop producing too.
  24. 4 min later
  25. 1:46:59 · watch on youtube.com · Dylan Patel

    So there is an angle of, the US' actions, from the angle of the expert controls, have been so inflammatory at slowing down China's progress on the leading edge that they've turned around and have accelerated their progress elsewhere because they know that this is so important.
  26. 1:47:20 · watch on youtube.com · Dylan Patel

    And so going back, can the US build it here? Yes, but it's going to take a ton of money. I truly think to revolutionize and completely in-source semiconductors would take a decade and a trillion dollars.
  27. 5 min later
  28. 1:52:35 · watch on youtube.com · Dylan Patel

    It's an objective fact that the world has been the most peaceful it's ever been when there are global hegemons, or regional hegemons in historical context. The Mediterranean was the most peaceful ever when the Romans were there.
  29. 6 min later
  30. 1:58:50 · watch on youtube.com · Dylan Patel

    FLOP is the vector that the government has cared about historically, but the other two vectors are arguably just as important. And especially when we come to this new paradigm, which the world is only just learning about over the last six months: reasoning.
  31. 15 min later
  32. 2:13:23 · watch on youtube.com · Dylan Patel

    OpenAI has a fantastic margin. When they're doing inference, their gross margins are north of 75%. So that's a four to five X factor right there of the cost difference, is that OpenAI is just making crazy amounts of money because they're the only one with the capability.
  33. 4 min later
  34. 2:17:46 · watch on youtube.com · Nathan Lambert

    We know that a lot of the American companies are very invested in safety, and that is the central culture of a place like Anthropic. And I think Anthropic sounds like a wonderful place to work, but if safety is your number one goal, it takes way longer to get artifacts out.
  35. 4 min later
  36. 2:21:25 · watch on youtube.com · Nathan Lambert

    these open models are probably going to keep coming for the time being, whether or not we want to stop them, and stopping them might make it even worse and harder to prepare.
  37. 7 min later
  38. 2:27:58 · watch on youtube.com · Dylan Patel

    There's this very good quote from Sam Altman who... He can be a hyperbeast sometimes, but one of the things he said, and I think I agree, is that superhuman persuasion will happen before superhuman intelligence, right? And if that's the case, then these things before we get this AGI ASI stuff, we can embed superhuman persuasion towards our ideal or whatever the ideal of the model maker is, right?
  39. 6 min later
  40. 2:34:18 · watch on youtube.com · Nathan Lambert

    I almost think it's practically impossible because you effectively have to remove them from the internet.
  41. 5 min later
  42. 2:39:47 · watch on youtube.com · Nathan Lambert

    And the important thing to say is that no matter how you want the model to behave, these RLHF and preference-tuning techniques also improve performance. So, on things like math evals and code evals, there is something innate to these, what is called contrastive loss functions.
  43. 4 min later
  44. 2:43:31 · watch on youtube.com · Nathan Lambert

    And these reasoning behaviors emerge naturally. So these things like, "Wait, let me see. Wait, let me check this. Oh, that might be a mistake." And they emerge from only having questions and answers.
  45. 5 min later
  46. 2:48:43 · watch on youtube.com · Dylan Patel

    I think it's actually probably simpler than that. It's probably something related to computer use or robotics rather than science discovery.
  47. 23 min later
  48. 3:11:27 · watch on youtube.com · Dylan Patel

    The important thing about, hey, is cost a limiting factor here? My view is that we'll have really awesome intelligence, like AGI, before we have it permeate throughout the economy.
  49. 6 min later
  50. 3:17:11 · watch on youtube.com · Dylan Patel

    But the funniest thing I think that comes out of this is Jevons paradox is true. AWS pricing for H100s has gone up over the last couple of weeks, since a little bit after Christmas, since V3 was launched, AWS H100 pricing has gone up.
  51. 2 min later
  52. 3:18:53 · watch on youtube.com · Nathan Lambert

    The more progress that AI makes or the higher the derivative of AI progress is, especially because NVIDIA's in the best place, the higher the derivative is, the sooner the market's going to be bigger and expanding and NVIDIA's the only one that does everything reliably right now.
  53. 1 min later
  54. 3:19:47 · watch on youtube.com · Dylan Patel

    One is ByteDance, arguably is the largest smuggler of GPUs for China. China's not supposed to have GPUs. ByteDance has over 500,000 GPUs. Why? Because they're all rented from companies around the world.
  55. 6 min later
  56. 3:26:04 · watch on youtube.com · Nathan Lambert

    Distillation is standard practice in industry. Whether or not, if you're at a closed lab where you care about terms of service and IP closely, you distill from your own models.
  57. 4 min later
  58. 3:30:29 · watch on youtube.com · Nathan Lambert

    I think that they're trying to shift the narrative. They're trying to protect themselves. We saw this years ago when ByteDance was actually banned from some OpenAI APIs for training on outputs. There's other AI startups that most people, if you're in the AI culture, were like they just told us they trained on OpenAI outputs and they never got banned.
  59. 2 min later
  60. 3:32:10 · watch on youtube.com · Dylan Patel

    So, Japan has a law which you're allowed to train on any training data and copyrights don't apply if you want to train a model, A. B, Japan has 9 gigawatts of curtailed nuclear power. C, Japan is allowed under the AI diffusion rule to import as many GPUs as they'd like.
  61. 2 min later
  62. 3:33:55 · watch on youtube.com · Nathan Lambert

    Code and data is hard, but ideas is easy. Silicon Valley operates on the way that top employees get bought out by other companies for a pay raise, and a large reason why these companies do this is to bring ideas with them.
  63. 12 min later
  64. 3:45:47 · watch on youtube.com · Dylan Patel

    Interesting thing is certain regions of the US transmitting power cost more than actually generating it because the grid is so slow to build. And the demand for power, and the ability to build power, and re-ramping on a natural gas plant or even a coal plant is easy enough to do, but transmitting the power's really hard.
  65. 19 min later
  66. 4:04:27 · watch on youtube.com · Dylan Patel

    But Google has never had that DNA of like, "This is a product we should sell." The Google Cloud, which is a separate organization from the TPU team, which is a separate organization from the DeepMind team, which is a separate organization from the Search team. There's a lot of bureaucracy here.
  67. 4 min later
  68. 4:08:39 · watch on youtube.com · Dylan Patel

    And they're decent, their hardware is better in many ways than in NVIDIA's. The problem is their software is really bad and I think they're getting better, right? They're getting better, faster, but the gulf is so large and they don't spend enough resources on it or haven't historically, right?
  69. 11 min later
  70. 4:19:37 · watch on youtube.com · Nathan Lambert

    The short-term that company that could make the most money is the one that figures out what advertising targeting method works for language model generations.
  71. 10 min later
  72. 4:29:46 · watch on youtube.com · Nathan Lambert

    And the history of NLP and language processing instruction, tuning and tasks per language model used to be like one language model did one task, and then in the instruction tuning literature, there's this point where you start adding more and more tasks together where it just starts to generalize to every task. And we don't know where on this curve we are.
  73. 2 min later
  74. 4:31:34 · watch on youtube.com · Dylan Patel

    But really the software engineering agents I think can be done faster sooner than any other agent because it is a verifiable domain. You can always unit test or compile, and there's many different regions of it can inspect the whole code base at once, which no engineer really can.
  75. 1 min later
  76. 4:32:38 · watch on youtube.com · Dylan Patel

    But what happens when every company can just invent their own business logic really cheaply and quickly? You stop using platform SaaS, you start building custom tailored solutions, you change them really quickly.
  77. 1 min later
  78. 4:33:47 · watch on youtube.com · Nathan Lambert

    The big picture is that I don't think it's going to be a cliff. I think a really good example of how growth changes is when Meta added stories. So Snapchat was on an exponential, they added stories, it flatlined.
  79. 2 min later
  80. 4:35:32 · watch on youtube.com · Nathan Lambert

    And humans are actually very good at reading or judging between two things versus... This goes back to the core of what RLHF and preference tuning is that it's hard to generate a good answer for a lot of problems, but it's easy to see which one is better.
  81. 11 min later
  82. 4:46:12 · watch on youtube.com · Nathan Lambert

    until there are feedback loops of open source AI, it seems like mostly an ideological mission. People like Mark Zuckerberg, which is like America needs this and I agree with him, but in the time where the motivation ideologically is high, we need to capitalize and build this ecosystem around, what benefits do you get from seeing the language model data?
  83. 17 min later
  84. 5:02:49 · watch on youtube.com · Nathan Lambert

    And for that reason, there's physical constraints to things like AGI, like recursive improvement to kill us all type stuff. For the physical reasons and for how humans have figured things out before, I'm not too worried about AI takeover.
  85. 1 min later
  86. 5:03:38 · watch on youtube.com · Dylan Patel

    it won't be one person rule them all, but it will be, the thing I worry about is it'll be few people, hundreds, thousands, tens of thousands, maybe millions of people rule whoever's left and the economy around it.