What Noam Brown thinks about scaling laws
Research scientist at OpenAI, who built the poker AIs Libratus and Pluribus and helped pioneer test-time compute and reasoning in language models.
Noam Brown did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
6 dated positions, 2026, in their own words. Our reading of what Noam Brown has said — not written or endorsed by them.
-
Their wordsThe problem is we're in a world now where the capability of the model is a function of how much money you put into it. Basically, if you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. Give it a budget of $10 million, you can do even more. At what budget should you evaluate these models? The policies that exist today don't really address that question.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 1st of 18 in this recording
-
Their wordsI think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented, the benchmark results are being presented in the wrong way. They're not controlling for the amount of test time compute that is being used on that benchmark question.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 2nd of 18 in this recording
-
Their wordsmy claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever or you plot the performance as a function of the amount of test time compute that's going into the model
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 4th of 18 in this recording
-
Their wordsthe preparedness frameworks and responsible scaling policies, they don't really account for the amount of test time compute.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 7th of 18 in this recording
-
Their wordsif you ask a person when was Abraham Lincoln born and they don't know the date. They could sit there, they could think about it for a week, if they if they don't have access to Wikipedia or something, they're not going to be able to do better answering that question if they thought about it for a week compared to 5 seconds. Same with the model.
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 12th of 18 in this recording
-
Their wordsand I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test time compute in order to achieve um their greatest intelligence. If you if it requires so much test time on compute to unlock the full capabilities of the model, then that means you're bottlenecked by time
↗Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brownyoutube.com 15th of 18 in this recording