Nathan Lambert, Dylan Patel did not write this page.
Every claim below is a statement made in this
recording, quoted word for word and linked to the second it was said, so
you can hear it rather than take our word for it. The wording comes from
the transcript published alongside the recording; the sentence above each
quote is our reading of the claim, not their wording.
Their wordsAnd we'll get into the details of the models and again and again as we try to get deeper into how the models were trained, we will say things like the data processing, data filtering data quality is the number one determinant of the model quality.
Their wordsThe scale word gets a lot of attention in this. The interpretation that I use is effectively to avoid adding the human priors to your learning process. And if you read the original essay, this is what it talks about is how researchers will try to come up with clever solutions to their specific problem that might get them small gains in the short term while simply enabling these deep learning systems to work efficiently, and for these bigger problems in the long term might be more likely to scale and continue to drive success.
Their wordsThis is why you want to work in post-training because the GPU cost for training is lower. So you can make a higher percentage of your training runs YOLO runs.
Their wordsI think it's even more impressive what OpenAI did in 2022. At the time, no one believed in mixture of experts models at Google who had all the researchers. OpenAI had such little compute and they devoted all of their compute for many months, all of it, 100% for many months to GPT-4 with a brand-new architecture with no belief that, "Hey, let me spend a couple of hundred million dollars, which is all of the money I have on this model." That is truly YOLO.
Their wordsAccepted practice is that for any given model that is a notable advancement, you're going to do two to 4x compute of the full training run in experiments alone.
Their wordsThere's not many worlds where China cannot train AI models. I think export controls are decapping the amount of compute or the density of compute that China can have.
Their wordsTo some extent, training a model does effectively nothing. They have a model. The thing that Dario is sort of speaking to is the implementation of that model, once trained to then create huge economic growth, huge increases in military capabilities, huge increases in productivity of people, betterment of lives.
Their wordsI think my personal definition of AGI is much simpler. I think language models are a form of AGI and all of this super powerful stuff is a next step that's great if we get these tools. But a language model has so much value in so many domains that it's a general intelligence to me.
Their wordsThere's some research that shows that the distribution is actually the limiting factor. So language models haven't yet made misinformation particularly change the equation there.
Their wordsif you believe we're in this sort of stage of economic growth and change that we've been in for the last 20 years, the export controls are absolutely guaranteeing that China will win long-term.
Their wordsChina, if they wanted to build the largest data center in the world, if they had access to the chips, could. So it's just a question of when, not if.
Their wordsSo there is an angle of, the US' actions, from the angle of the expert controls, have been so inflammatory at slowing down China's progress on the leading edge that they've turned around and have accelerated their progress elsewhere because they know that this is so important.
Their wordsAnd so going back, can the US build it here? Yes, but it's going to take a ton of money. I truly think to revolutionize and completely in-source semiconductors would take a decade and a trillion dollars.
Their wordsIt's an objective fact that the world has been the most peaceful it's ever been when there are global hegemons, or regional hegemons in historical context. The Mediterranean was the most peaceful ever when the Romans were there.
Their wordsFLOP is the vector that the government has cared about historically, but the other two vectors are arguably just as important. And especially when we come to this new paradigm, which the world is only just learning about over the last six months: reasoning.
Their wordsOpenAI has a fantastic margin. When they're doing inference, their gross margins are north of 75%. So that's a four to five X factor right there of the cost difference, is that OpenAI is just making crazy amounts of money because they're the only one with the capability.
Their wordsWe know that a lot of the American companies are very invested in safety, and that is the central culture of a place like Anthropic. And I think Anthropic sounds like a wonderful place to work, but if safety is your number one goal, it takes way longer to get artifacts out.
Their wordsthese open models are probably going to keep coming for the time being, whether or not we want to stop them, and stopping them might make it even worse and harder to prepare.
Their wordsThere's this very good quote from Sam Altman who... He can be a hyperbeast sometimes, but one of the things he said, and I think I agree, is that superhuman persuasion will happen before superhuman intelligence, right? And if that's the case, then these things before we get this AGI ASI stuff, we can embed superhuman persuasion towards our ideal or whatever the ideal of the model maker is, right?
Their wordsAnd the important thing to say is that no matter how you want the model to behave, these RLHF and preference-tuning techniques also improve performance. So, on things like math evals and code evals, there is something innate to these, what is called contrastive loss functions.
Their wordsAnd these reasoning behaviors emerge naturally. So these things like, "Wait, let me see. Wait, let me check this. Oh, that might be a mistake." And they emerge from only having questions and answers.
Their wordsI think it's actually probably simpler than that. It's probably something related to computer use or robotics rather than science discovery.
Their wordsThe important thing about, hey, is cost a limiting factor here? My view is that we'll have really awesome intelligence, like AGI, before we have it permeate throughout the economy.
Their wordsBut the funniest thing I think that comes out of this is Jevons paradox is true. AWS pricing for H100s has gone up over the last couple of weeks, since a little bit after Christmas, since V3 was launched, AWS H100 pricing has gone up.
Their wordsThe more progress that AI makes or the higher the derivative of AI progress is, especially because NVIDIA's in the best place, the higher the derivative is, the sooner the market's going to be bigger and expanding and NVIDIA's the only one that does everything reliably right now.
Their wordsOne is ByteDance, arguably is the largest smuggler of GPUs for China. China's not supposed to have GPUs. ByteDance has over 500,000 GPUs. Why? Because they're all rented from companies around the world.
Their wordsDistillation is standard practice in industry. Whether or not, if you're at a closed lab where you care about terms of service and IP closely, you distill from your own models.
Their wordsI think that they're trying to shift the narrative. They're trying to protect themselves. We saw this years ago when ByteDance was actually banned from some OpenAI APIs for training on outputs. There's other AI startups that most people, if you're in the AI culture, were like they just told us they trained on OpenAI outputs and they never got banned.
Their wordsSo, Japan has a law which you're allowed to train on any training data and copyrights don't apply if you want to train a model, A. B, Japan has 9 gigawatts of curtailed nuclear power. C, Japan is allowed under the AI diffusion rule to import as many GPUs as they'd like.
Their wordsCode and data is hard, but ideas is easy. Silicon Valley operates on the way that top employees get bought out by other companies for a pay raise, and a large reason why these companies do this is to bring ideas with them.
Their wordsInteresting thing is certain regions of the US transmitting power cost more than actually generating it because the grid is so slow to build. And the demand for power, and the ability to build power, and re-ramping on a natural gas plant or even a coal plant is easy enough to do, but transmitting the power's really hard.
Their wordsBut Google has never had that DNA of like, "This is a product we should sell." The Google Cloud, which is a separate organization from the TPU team, which is a separate organization from the DeepMind team, which is a separate organization from the Search team. There's a lot of bureaucracy here.
Their wordsAnd they're decent, their hardware is better in many ways than in NVIDIA's. The problem is their software is really bad and I think they're getting better, right? They're getting better, faster, but the gulf is so large and they don't spend enough resources on it or haven't historically, right?
Their wordsThe short-term that company that could make the most money is the one that figures out what advertising targeting method works for language model generations.
Their wordsAnd the history of NLP and language processing instruction, tuning and tasks per language model used to be like one language model did one task, and then in the instruction tuning literature, there's this point where you start adding more and more tasks together where it just starts to generalize to every task. And we don't know where on this curve we are.
Their wordsBut really the software engineering agents I think can be done faster sooner than any other agent because it is a verifiable domain. You can always unit test or compile, and there's many different regions of it can inspect the whole code base at once, which no engineer really can.
Their wordsBut what happens when every company can just invent their own business logic really cheaply and quickly? You stop using platform SaaS, you start building custom tailored solutions, you change them really quickly.
Their wordsThe big picture is that I don't think it's going to be a cliff. I think a really good example of how growth changes is when Meta added stories. So Snapchat was on an exponential, they added stories, it flatlined.
Their wordsAnd humans are actually very good at reading or judging between two things versus... This goes back to the core of what RLHF and preference tuning is that it's hard to generate a good answer for a lot of problems, but it's easy to see which one is better.
Their wordsuntil there are feedback loops of open source AI, it seems like mostly an ideological mission. People like Mark Zuckerberg, which is like America needs this and I agree with him, but in the time where the motivation ideologically is high, we need to capitalize and build this ecosystem around, what benefits do you get from seeing the language model data?
Their wordsAnd for that reason, there's physical constraints to things like AGI, like recursive improvement to kill us all type stuff. For the physical reasons and for how humans have figured things out before, I'm not too worried about AI takeover.
Their wordsit won't be one person rule them all, but it will be, the thing I worry about is it'll be few people, hundreds, thousands, tens of thousands, maybe millions of people rule whoever's left and the economy around it.