korrents

korrents · Y Combinator

Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work

Dmitri Dolgov · 49m · youtube.com

22 korrents from this recording

Dmitri Dolgov did not write this page.

Every claim below is a statement made in this recording, quoted word for word and linked to the second it was said, so you can hear it rather than take our word for it. The wording comes from the transcript published alongside the recording; the sentence above each quote is our reading of the claim, not their wording.

  1. 2:11 · watch on youtube.com

    That the best AI moments will look like nothing happened. It's just the task got done safely and smoothly.
  2. 1 min later
  3. 3:07 · watch on youtube.com

    in Silicon Valley, there's a common mantra to move fast and break things. However, when you're dealing with atoms instead of bits, breaking things is not really okay. So, the thing you have to do is to move fast and ship safely.
  4. 1 min later
  5. 4:11 · watch on youtube.com

    In the physical world, the cost of a mistake can be measured in human lives, not tokens. There's simply not an undo and a retry button.
  6. 2 min later
  7. 5:43 · watch on youtube.com

    Given the high cost of errors, you need to have a very high level of safety and a very high level of confidence on day one before you deploy your first robot, before you drive your your first autonomous mile.
  8. 3 min later
  9. 8:21 · watch on youtube.com

    And a working demo is 1% at best of the work that you have to do. The many nines of performance, the many nines of reliability that follow, that's where the real work happens.
  10. 2 min later
  11. 10:30 · watch on youtube.com

    It took us about 10 more years to begin providing a service, and then 5 more years to scale to half a million trips per week. So, the demo took 18 months, the product took about 15 years.
  12. 1 min later
  13. 11:44 · watch on youtube.com

    reliability and performance lives on this exponential ladder of nines. So, getting to that first 90% or 99%, that's the easy part. But then every next nine that you want to add, that takes about 10 times more effort.
  14. 1 min later
  15. 12:21 · watch on youtube.com

    And at scale, the long tail is the problem space, is your entire problem statement. When you drive millions of miles per week, a rare event that might happen once in a million miles, that just becomes your daily reality.
  16. 2 min later
  17. 14:01 · watch on youtube.com

    And that's why every hype cycle produces a wave of absolutely spectacular demos and very few real products. And the recurring mistake of every cycle is spending on the demo when you should be saving for the nines.
  18. 1 min later
  19. 15:16 · watch on youtube.com

    common failure mode is picking the tech that gives you the fastest early ramp, riding that steep curve, feeling like you're winning, projecting that, you know, steep slope into the future and feeling like the sky is the limit, and then hitting the plateau, and discovering that the technology path that you picked actually flattens out way before the performance that is required by your product.
  20. 1 min later
  21. 16:37 · watch on youtube.com

    However, if you are targeting full autonomy, and you're targeting superhuman, strongly superhuman performance, you find that weak sensing just leads to a safety curve that flattens out way too early.
  22. 4 min later
  23. 20:41 · watch on youtube.com

    So, betting your company, betting your approach on today's hardware prices is just betting your company on a number that has a fairly short shelf life and is going to expire.
  24. 1 min later
  25. 22:03 · watch on youtube.com

    Turns out, the task of driving is not that dissimilar from the task of modeling language uh because of the social aspects of driving. You're kind of having a conversation with other dynamic actors in the world, but you're doing that in the space kind of body language of your agent, your car, as opposed to just the language of words.
  26. 9 min later
  27. 31:25 · watch on youtube.com

    Essentially, structure that fights scale will always lose. And structure that channels scale always wins.
  28. 1 min later
  29. 32:11 · watch on youtube.com

    But if you need to reach superhuman levels of performance in a fully autonomous agent in a safety-critical environment, uh, just doing kind of that basic vanilla end-to-end is not enough.
  30. 4 min later
  31. 36:18 · watch on youtube.com

    So, the lesson here is to bet on a system that's maximally learned and minimally constrained and leverage structure intentionally to boost performance and scaling laws both in training and in evaluation.
  32. 2 min later
  33. 38:06 · watch on youtube.com

    And the problem of building a good realistic simulator is just as hard as building the agent itself.
  34. 4 min later
  35. 42:20 · watch on youtube.com

    So, a deployment of your agent uh in the real world generates data. That data then grounds the simulator and makes it more realistic. The simulator generates harder edge cases for the critic to score and for the agent to learn from.
  36. 1 min later
  37. 42:58 · watch on youtube.com

    that your model is really table stakes, but eval and metrics, that's your most important. That's your strategic moat. So, build your eval before you build your technology.
  38. 43:16 · watch on youtube.com

    If you can't quantitatively define what good enough means, you're not really building a product, you're just iterating on your demo.
  39. 2 min later
  40. 45:37 · watch on youtube.com

    Your your models can be leaked, algorithms can be replicated, but hundreds of millions of miles of fully autonomous operations in the real world, backed by evidence-grade evaluation and publicly audited proof, that is much, much more difficult to replicate.
  41. 1 min later
  42. 46:49 · watch on youtube.com

    in the areas where we operate, the Waymo driver is about 17 times better than human drivers when it comes to crashes uh that cause serious injury.