Every claim below is a statement made in this
recording, quoted word for word and linked to the second it was said, so
you can hear it rather than take our word for it. The wording comes from
the transcript published alongside the recording; the sentence above each
quote is our reading of the claim, not their wording.
Their wordsSo like they don't have enough intelligence. They're not multimodal enough. They can't do computer use and all this kind of stuff. And they don't do a lot of things. You know, they don't have continual learning. You can't just tell them something and they'll remember it. And they're just cognitively lacking. And it's just not working. And I just think that it will take about a decade to work through all of those issues.
Their wordsAnd it just so turns out that um this was extremely early, way too early. so early that we shouldn't have been working on that, you know, uh because um if you're just stumbling your way around and keyboard mashing and mouse clicking and trying to get rewards in these environments, um your reward is too sparse and you just won't learn and you're going to burn a forest uh computing and you're never actually going to get something off the ground.
Their wordswe're not doing training by evolution. Uh we're doing training by basically imitation of humans and the data that they've put on the internet. And so you end up with these like sort of ethereal spirit entities because they're fully digital and they're kind of like mimicking humans.
Their wordsa lot of what looks like learning is actually a lot more maturation of the brain and I think that actually very little reinforcement learning for animals and I think a lot of the reinforcement learning is actually like more like motor tasks. It's not intelligence tasks. So I actually kind of think humans don't actually like really use RL roughly speaking is what I would say.
Their wordsSo that's why I kind of call pre-training this kind of like crappy evolution. It's like the practically possible version with our technology and what we have available to us to get to a starting point where we can actually do things like reinforcement learning and so on.
Their wordsAnd actually, you don't actually need or want the knowledge. I actually think that's probably actually holding back the neural networks overall because it's actually like getting them to rely on the knowledge a little too much sometimes.
Their wordsanything that's in the weights it's kind of like a hazy recollection of what you read a year ago anything that you give it as a context uh at test time is directly in the working memory
Their wordsThese models don't really have this distillation phase um of taking what happened, analyzing it, obsessively thinking through it, um basically doing some kind of a synthetic data generation process and distilling it back back into the weights
Their wordsSo, don't write blog posts, don't do slides, don't do any of that. Like, build the code, arrange it, get it to work. It's the only way to go. Otherwise, you're missing knowledge.
Their wordsI think um yeah, they're not very good at code that hasn't never been written before maybe is like one way to put it, which is like what we're trying to achieve when we're building these models.
Their wordsAnd I kind of feel like the industry it's it's um it's over it's it's making too big of a jump and it's trying to pretend like this is amazing and it's not. It's slop and I think they're not coming to terms with it and maybe they're trying to fund raise or something like that.
Their wordsreinforcement learning is a lot worse than I think the average person thinks reinforcement learning is terrible. It just so happens that uh everything that we had before is much worse
Their wordsA human would never do this. Number one, a human would never do hundreds of rollouts. Uh number two, when a person sort of finds a solution, they will have a pretty complicated process of review of like, okay, I think these parts that I did well, these parts I did not do that well, I should probably do this or that.
Their wordsthe reason that I think this is kind of tricky is quite subtle. And it's the fact that anytime you use an LLM to assign a reward, those LLMs are giant things with billions of parameters and they're gameable.
Their wordsall of the samples you get from models are silently collapsed. They're silently, this is not obvious if you look at any individual example of it. They occupy a very tiny manifold of the possible space of um sort of thoughts about content.
Their wordsI also think humans collapse over time. Uh I think this is uh again these analogies are surprisingly good but humans collapse during the course of their lives. This is why children have completely u you know they haven't overfit yet and they will say stuff that will shock you
Their wordsUm so I almost feel like because the internet is so terrible, we actually have to sort of like build really big models to compress all that. Uh most of that compression is memory work instead of like cognitive work. But what we really want is the cognitive part actually delete the memory
Their wordsNow, number one, the first concession that people make all the time is they just take out all the physical stuff because we're just talking about digital knowledge work. I feel like that's a pretty major concession compared to the original definition which was like any task a human can do.
Their wordsso a good example recently was um Jeff Hinton's prediction that radiologists would not be a job anymore and this turned out to be very wrong in a bunch of ways right so radiologists are alive and well and growing even though computer vision is really really good at recognizing all the different things that they have to recognize
Their wordsBut even there I'm not actually looking at full automation yet. I'm looking for an autonomy slider and I almost expect that we are not going to instantly replace people. We're going to be swapping in AIs that do 80% of the volume. They delegate 20% of the volume to humans and humans are supervising teams of five AIs doing the call center work that's more rote.
Their wordsso I think there's there's an interesting point here because I do believe coding is like the perfect first thing for uh for a for uh these LLMs and uh agents and that's because coding has always fundamentally uh worked around text.
Their wordsand we'll gradually layer all this stuff everywhere and there will be fewer and fewer people who understand it and that there will be a sort of this like scenario of a gradual loss of control and understanding of what's happening that to me seems most likely outcome of how all this stuff will go down.
Their wordsI do, but it's business as usual because we're we're in an intelligence explosion already and have been for decades. And when you look at GDP, it's basically the GDP curve that is an exponential weighted sum over so many aspects of the industry. Everything is gradually being automated has been for hundreds of years.
Their wordsBut I guess like the notion of culture and of written record and of like passing down notes between each other. I don't think there's an equivalent of that with LM right now. So LM don't really have culture right now and it's kind of like one of the I think uh impediments I would say.
Their wordsI think basically what takes the long amount of time and the way to think about it is that it's a march of nines and every single nine is a constant amount of work.
Their wordsI think some of the times they are but they're certainly involved and there are people and in some sense we haven't actually removed the person we've like moved them to somewhere where we can't see them.