Every claim below is a statement made in this
recording, quoted word for word and linked to the second it was said, so
you can hear it rather than take our word for it. The wording comes from
the transcript published alongside the recording; the sentence above each
quote is our reading of the claim, not their wording.
Their wordsSo to me like what what I tend to think about a lot in terms of timelines is not the date when it will be done but the date when it will when like the flywheel starts basically.
Their wordsLike if you answer a question, you just like answered it wrong. It's like well it's not like you can just like go back and like tweak a few things like the person you told the answer to might not even know that it's wrong. Whereas if you're like folding the t-shirt and you messed up a little bit like it's pretty obvious like you can reflect on that, figure out what happened and do it better next time.
Their wordsSo, I think it'll be the same thing that that we'll see an increase in the scope that we're giving that we're willing to give to the robots as they get better and better where initially the scope might be like there is a particular thing you do like you're making the coffee or something. Uh whereas as they get more capable, as their ability to have common sense and a broader repertoire of tasks increases, then we'll give them greater scope. Now you're running the whole coffee shop.
Their wordsUh but it uh I think we kind of like know like roughly the puzzle pieces and it's something that we need to work on and I think if we work on it and we're a bit lucky and everything kind of goes as planned I think single digit is reasonable.
Their wordsAnd I think we'll see the same thing with automation where uh basically robot plus human is much better than just human or just robot. Uh and and that just like makes total sense. It also makes it much easier to get all the technology bootstrapped because when it's robot plus human, now there's a lot more potential for the robot to like actually learn on the job, acquire new skills.
Their wordsLike now, like basically learning is not for these systems is not just learning from raw actions. It's also learning from words eventually be learning from observing what people do from the kind of natural feedback that you receive when you're doing a job together with somebody else.
Their wordsSo, that's not an argument about robotics being easier than autonomous driving. It's just an argument for 2025 being a better year than 2009.
Their wordsAnd when you make a mistake and correct it, well, first you you've achieved the task because you've corrected, but you've also gained knowledge that allows you to avoid that mistake in the future. With driving, because of the dynamics of how it's set up, it's very hard to make a mistake, correct it, and then learn from it because the mistakes themselves have significant ramifications.
Their wordsuh common sense meaning the ability to make inferences about what might happen uh that are reasonable guesses but that do not require you to experience that mistake and learn and learn from it in advance that's tremendously important and that's something that we basically had no idea how to do uh about 5 years ago but now uh you we can actually use LLMs and VLMs ask them questions and they will make reasonable guesses
Their wordsbut um to make robotic foundation models really work it's not just a laboratory science uh kind of experiment. It's also uh it also requires kind of industrial scale uh building effort like it's it's like it's more like the Apollo program than it is like a science experiment
Their wordsBut because we don't know the answer to that, to me, a much more useful way to think about it is not how much data do we need to get before we're fully done, but how much data do we need to get before we can get started, meaning before we can get uh a data flywheel that represents a self- sustaining uh and ever growing data collection.
Their wordsWhereas with text, it's already sort of been abstract into those bits that we as humans care about. So the representations are already there and they're not just good representations. They actually like focus in on what really matters.
Their wordsUh and its perception is in service to fulfilling that purpose. And that is like a really great uh focusing factor. We know that for people this really matters. Like literally what you see is affected by what you're trying to do.
Their wordsYeah. So there's a subtlety here. Emerging capabilities don't just come from the fact that internet data has a lot of stuff in it. They also come from the fact that generalization once it reaches a certain level becomes compositional.
Their wordsBut the reason why it's not the most important thing for the kind of skills that you saw when you visited us, it at some level I think it comes back to Moravik's paradox. So Morovik's paradox is basically that it's like you know if you know one thing about if you want to know one thing about robotics it's like that's that's the thing. Morovik's paradox says that basically uh in AI the easy things are hard and the hard things are easy.
Their wordsit's something like this that the brain is extremely parallel. uh it kind of has to be just out of just because of the biohysics. U but like it's even more parallel than your GPU.
Their wordsUh so in order to effectively learn from your own experience, it turns out that it's really really important to already know something about what you're doing. Otherwise, it takes far too long.
Their wordsI think that it's optimistically that it's actually the other way around that the robotics uh element of the equation will make all the other stuff better. And there are two uh reasons for this that that I could tell you about. One has to do with representations and focus.
Their wordsbut something to remember is that when a pilot is using a simulator to learn to fly an airplane, they're extremely goal- directed. So, their goal in life is not to learn to use a simulator. Their goal in life is to learn to fly the airplane.
Their wordsSo here's what I would say that deep down at a very fundamental level the synthetic experience that you create yourself doesn't allow you to learn more about the world. It allows you to rehearse things. It allows you to consider counterfactuals but somehow information about the world needs to get injected into the system.
Their wordsLike people are people and robots are robots. Like the the better analogy for the robot, it's it's like your car or a bulldozer. Uh like uh it has much lower maintenance requirements. You can put them into all sorts of weird places and they don't have to look like people at all.
Their wordsSo traditional robots and factories uh they need to make motions that are highly repeatable and therefore it requires a degree of precision and robustness that you don't need if you can use cheap visual feedback. So AI also makes robots more affordable uh and lowers the requirements on the hardware.
Their wordsrobots help with uh physical things uh physical work. And if producing robots is itself physical work, then getting really good at robotics should help with that. It's a little circular, of course.