korrents

What Noam Brown thinks about benchmarks

@noam-brown · 18 positions · 0 changes of mind

Research scientist at OpenAI, who built the poker AIs Libratus and Pluribus and helped pioneer test-time compute and reasoning in language models.

Noam Brown did not write this page.

We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.

8 dated positions, 2026, in their own words. Our reading of what Noam Brown has said — not written or endorsed by them.

  1. The problem is we're in a world now where the capability of the model is a function of how much money you put into it. Basically, if you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. Give it a budget of $10 million, you can do even more. At what budget should you evaluate these models? The policies that exist today don't really address that question.

    Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 1st of 18 in this recording

  2. I think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented, the benchmark results are being presented in the wrong way. They're not controlling for the amount of test time compute that is being used on that benchmark question.

    Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 2nd of 18 in this recording

    scaling laws

  3. But what we're seeing today with the modern models is that 5.5 and other models can think for if you scaffold them reasonably well, can think for weeks even um before having performance plateau on some of these benchmarks. And so, the point at which they plateau is simply too far out to reasonably test.

    Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 3rd of 18 in this recording

  4. my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model

    Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 4th of 18 in this recording

    scaling laws

  5. Um so if you say okay well we're going to instead of just running this model once we're going to run it five times and take the best of the five responses or like ask a judge which one it thinks is best then you can get much higher scores than that model. And so it's really easy to make something that looks a lot better on paper but is actually not better once you control for the amount of test time compute.

    Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 5th of 18 in this recording

  6. And the problem is if you want to evaluate the capabilities of a model, what it can do after running for a month, the only way to be fully sure is to actually run it for a month. And if you want to know after 6 months, the only way to know fully is to run it for six months.

    Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 8th of 18 in this recording

  7. It's actually very difficult because, yeah, you would have to the only way to to really do the evaluations is then delay the model release cycle. Um and you know there's a lot of competitive pressure right now to not do that.

    Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 9th of 18 in this recording

  8. And so you kind of end up in this this bad equilibrium where everybody kind of knows that it's a bad equilibrium, but like nobody wants to break out. And I I felt like, okay, well, if I just hopefully come out and say like, look guys, let's all recognize that we're in a bad equilibrium and let's move to this different equilibrium where we're we're plotting things with an X-axis

    Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 18th of 18 in this recording