korrents

korrents · No Priors

Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown

Noam Brown · 36m · youtube.com

18 korrents from this recording

Noam Brown did not write this page.

Every claim below is a statement made in this recording, quoted word for word and linked to the second it was said, so you can hear it rather than take our word for it. The wording comes from the transcript published alongside the recording; the sentence above each quote is our reading of the claim, not their wording.

  1. 0:13 · watch on youtube.com

    The problem is we're in a world now where the capability of the model is a function of how much money you put into it. Basically, if you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. Give it a budget of $10 million, you can do even more. At what budget should you evaluate these models? The policies that exist today don't really address that question.
  2. 2 min later
  3. 2:27 · watch on youtube.com

    I think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented, the benchmark results are being presented in the wrong way. They're not controlling for the amount of test time compute that is being used on that benchmark question.
  4. 1 min later
  5. 3:31 · watch on youtube.com

    But what we're seeing today with the modern models is that 5.5 and other models can think for if you scaffold them reasonably well, can think for weeks even um before having performance plateau on some of these benchmarks. And so, the point at which they plateau is simply too far out to reasonably test.
  6. 1 min later
  7. 4:02 · watch on youtube.com

    my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model
  8. 3 min later
  9. 7:18 · watch on youtube.com

    Um so if you say okay well we're going to instead of just running this model once we're going to run it five times and take the best of the five responses or like ask a judge which one it thinks is best then you can get much higher scores than that model. And so it's really easy to make something that looks a lot better on paper but is actually not better once you control for the amount of test time compute.
  10. 4 min later
  11. 11:16 · watch on youtube.com

    And I wouldn't be surprised if you know 6 months or a year from now, the model is able to do zero shot an entire poker solver, basically my entire PhD thesis in one go.
  12. 2 min later
  13. 12:52 · watch on youtube.com

    the preparedness frameworks and responsible scaling policies, they don't really account for the amount of test time compute.
  14. 3 min later
  15. 15:54 · watch on youtube.com

    And the problem is if you want to evaluate the capabilities of a model, what it can do after running for a month, the only way to be fully sure is to actually run it for a month. And if you want to know after 6 months, the only way to know fully is to run it for six months.
  16. 1 min later
  17. 16:54 · watch on youtube.com

    It's actually very difficult because, yeah, you would have to the only way to to really do the evaluations is then delay the model release cycle. Um and you know there's a lot of competitive pressure right now to not do that.
  18. 2 min later
  19. 18:58 · watch on youtube.com

    Um but it would be possible and it would have been possible for somebody to disprove the erdos unit distance conjecture before we did using a general purpose model. And nobody had explored sufficiently what happens if I put $100,000 worth of compute into 5.5 what could it do?
  20. 1 min later
  21. 20:14 · watch on youtube.com

    we are trying to encourage people to not spend all their time just like going through all the mathematical open problems physics problems and um just seeing pushing the models to their limits to see what they can prove or disprove.
  22. 2 min later
  23. 21:54 · watch on youtube.com

    if you ask a person when was Abraham Lincoln born and they don't know the date. They could sit there, they could think about it for a week, if they if they don't have access to Wikipedia or something, they're not going to be able to do better answering that question if they thought about it for a week compared to 5 seconds. Same with the model.
  24. 1 min later
  25. 23:18 · watch on youtube.com

    one thing I see for research in particular is they don't have very good research taste right now and so I think they're actually a very good complement to researchers
  26. 1 min later
  27. 24:18 · watch on youtube.com

    And then I was like, okay, can you come up with an algorithm that is better than the algorithms that I came up with or that anybody else came up with and go ahead and like look at all the published work and synthesize that and then try to come up with something novel and it it's not able to do it. And I can give it a lot of time and it's it's still not able to do it.
  28. 2 min later
  29. 26:20 · watch on youtube.com

    and I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test time compute in order to achieve um their greatest intelligence. If you if it requires so much test time on compute to unlock the full capabilities of the model, then that means you're bottlenecked by time
  30. 2 min later
  31. 27:50 · watch on youtube.com

    it's not that humans have become smarter over it's not that they evolved to become smarter over, you know, the past 50,000 years. It's that humans are able to do a lot more today than they were back in caveman times because there have been billions of humans thinking for a long time and building off of each other's accumulated knowledge.
  32. 4 min later
  33. 31:38 · watch on youtube.com

    And I think they're at a point now where they've actually been at a point for a while now where I feel like I can just trust the outputs arguably more than I could trust the output from from a human,
  34. 1 min later
  35. 32:57 · watch on youtube.com

    And so you kind of end up in this this bad equilibrium where everybody kind of knows that it's a bad equilibrium, but like nobody wants to break out. And I I felt like, okay, well, if I just hopefully come out and say like, look guys, let's all recognize that we're in a bad equilibrium and let's move to this different equilibrium where we're we're plotting things with an X-axis