korrents

korrents · a16z Podcast

Why Fei-Fei Li Is Betting on Spatial Intelligence

Fei-Fei Li · 43m · youtube.com

13 korrents from this recording

Fei-Fei Li did not write this page.

Every claim below is a statement made in this recording, quoted word for word and linked to the second it was said, so you can hear it rather than take our word for it. The wording comes from the transcript published alongside the recording; the sentence above each quote is our reading of the claim, not their wording.

  1. 7:17 · watch on youtube.com

    It's the first time we have a unification of pixel generation and pixel reconstruction. In the world of computer vision this field has been around for more than half a century.
  2. 2 min later
  3. 8:50 · watch on youtube.com

    Well, spatial intelligence eventually must enable us to both generate what the space is, reason within it, and being able to edit and interact within it.
  4. 1 min later
  5. 9:20 · watch on youtube.com

    But to do that, a fundamental problem to solve is to understand the geometry and structure and the physics of the space.
  6. 9:30 · watch on youtube.com

    And I do believe Atlas is a significant step forward because now with every single frame, you have a you can generate an estimate a important piece of information, which is the the view viewpoint, the camera pose. And that is the most critical information one needs about the geometry of the of the space.
  7. 1 min later
  8. 10:03 · watch on youtube.com

    So in the on the path to spatial intelligence, generating pixels is definitely a a early step, which we have seen with what you call it gazillions of models. But generating pixels that are truly spatially contextualized and grounded is absolutely another major step. And that is the very hard step that Atlas has taken.
  9. 7 min later
  10. 16:55 · watch on youtube.com

    One thing that's under appreciated on the website of the demos is the Stanford demo where Ben showed uh anywhere between 3 to 25 images you can reconstruct that entire Stanford quad. But the thing is we had to show it from aerial view. But every single input image is been standing on the ground taking a picture from the ground. So, everything you see are generated but according to the laws of reconstruction.
  11. 5 min later
  12. 21:39 · watch on youtube.com

    I think three of us have total conviction about the scaling law. That that I think we do. I do think the exact architecture choices and data mixtures is where the the devils are in the details.
  13. 9 min later
  14. 30:49 · watch on youtube.com

    We should zoom out and recognize the biggest problem right now in robotics is actually data. One day it'll be chips, but for now it's data.
  15. 31:17 · watch on youtube.com

    There is also a very important step called randomization. Is that you have to take the same environment and then randomize the conditions. So, the cable doesn't literally only, you know, uh bend this way. It can bend a different way or the box can have different sizes, colors, different lids, and all that.
  16. 1 min later
  17. 32:02 · watch on youtube.com

    cuz we don't yet have a a frontier foundation model that's robust enough for for uh robotics.
  18. 9 min later
  19. 40:34 · watch on youtube.com

    I think for me, let's go back to the first principle of intelligence. Intelligence is not sitting there stuck and just seeing something or interpreting something when it comes to space and physical space, right? It's really this uh closing the loop between seeing and experiencing and interaction.
  20. 2 min later
  21. 42:40 · watch on youtube.com

    So okay, so to take a evolutionary view, right? That new viewpoint prediction is exactly evolution had to solve by making animals move. You you nature give animals eyes. But nature didn't give trees eye. Eyes. Why? Because when you move, you see a new viewpoint.
  22. 1 min later
  23. 43:10 · watch on youtube.com

    we do believe very strongly that next viewpoint prediction is is the equivalent of next token prediction.