Noam Brown
Research scientist at OpenAI who built the poker AIs Libratus and Pluribus and helped pioneer test-time compute and reasoning in language models.
Noam Brown did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
-
Their wordsThe problem is we're in a world now where the capability of the model is a function of how much money you put into it. Basically, if you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. Give it a budget of $10 million, you can do even more. At what budget should you evaluate these models? The policies that exist today don't really address that question.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 1st of 18 in this recording
-
Their wordsI think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented, the benchmark results are being presented in the wrong way. They're not controlling for the amount of test time compute that is being used on that benchmark question.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 2nd of 18 in this recording
-
Their wordsBut what we're seeing today with the modern models is that 5.5 and other models can think for if you scaffold them reasonably well, can think for weeks even um before having performance plateau on some of these benchmarks. And so, the point at which they plateau is simply too far out to reasonably test.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 3rd of 18 in this recording
-
Their wordsmy claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 4th of 18 in this recording
-
Their wordsUm so if you say okay well we're going to instead of just running this model once we're going to run it five times and take the best of the five responses or like ask a judge which one it thinks is best then you can get much higher scores than that model. And so it's really easy to make something that looks a lot better on paper but is actually not better once you control for the amount of test time compute.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 5th of 18 in this recording
-
Their wordsAnd I wouldn't be surprised if you know 6 months or a year from now, the model is able to do zero shot an entire poker solver, basically my entire PhD thesis in one go.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 6th of 18 in this recording
-
Their wordsthe preparedness frameworks and responsible scaling policies, they don't really account for the amount of test time compute.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 7th of 18 in this recording
-
Their wordsAnd the problem is if you want to evaluate the capabilities of a model, what it can do after running for a month, the only way to be fully sure is to actually run it for a month. And if you want to know after 6 months, the only way to know fully is to run it for six months.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 8th of 18 in this recording
-
Their wordsIt's actually very difficult because, yeah, you would have to the only way to to really do the evaluations is then delay the model release cycle. Um and you know there's a lot of competitive pressure right now to not do that.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 9th of 18 in this recording
-
Their wordsUm but it would be possible and it would have been possible for somebody to disprove the erdos unit distance conjecture before we did using a general purpose model. And nobody had explored sufficiently what happens if I put $100,000 worth of compute into 5.5 what could it do?
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 10th of 18 in this recording
-
Their wordswe are trying to encourage people to not spend all their time just like going through all the mathematical open problems physics problems and um just seeing pushing the models to their limits to see what they can prove or disprove.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 11th of 18 in this recording
-
Their wordsif you ask a person when was Abraham Lincoln born and they don't know the date. They could sit there, they could think about it for a week, if they if they don't have access to Wikipedia or something, they're not going to be able to do better answering that question if they thought about it for a week compared to 5 seconds. Same with the model.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 12th of 18 in this recording
-
Their wordsone thing I see for research in particular is they don't have very good research taste right now and so I think they're actually a very good complement to researchers
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 13th of 18 in this recording
-
Their wordsAnd then I was like, okay, can you come up with an algorithm that is better than the algorithms that I came up with or that anybody else came up with and go ahead and like look at all the published work and synthesize that and then try to come up with something novel and it it's not able to do it. And I can give it a lot of time and it's it's still not able to do it.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 14th of 18 in this recording
-
Their wordsand I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test time compute in order to achieve um their greatest intelligence. If you if it requires so much test time on compute to unlock the full capabilities of the model, then that means you're bottlenecked by time
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 15th of 18 in this recording
-
Their wordsit's not that humans have become smarter over it's not that they evolved to become smarter over, you know, the past 50,000 years. It's that humans are able to do a lot more today than they were back in caveman times because there have been billions of humans thinking for a long time and building off of each other's accumulated knowledge.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 16th of 18 in this recording
-
Our reading
On everyday high-stakes questions, a model's answer can now be trusted more than an expert human's.
Their wordsAnd I think they're at a point now where they've actually been at a point for a while now where I feel like I can just trust the outputs arguably more than I could trust the output from from a human,
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 17th of 18 in this recording
-
Their wordsAnd so you kind of end up in this this bad equilibrium where everybody kind of knows that it's a bad equilibrium, but like nobody wants to break out. And I I felt like, okay, well, if I just hopefully come out and say like, look guys, let's all recognize that we're in a bad equilibrium and let's move to this different equilibrium where we're we're plotting things with an X-axis
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 18th of 18 in this recording