benchmarks
FilterEveryone, all time
- DR
Dax Raad quoted
Their wordsevery week i see a new benchmark that ranks claude code last and everyone pats themselves on the back for not using claude code and being "smarter" and it's still #1 and still growing faster than most of the other things on the list this has been going on for a year
- DR
Dax Raad quoted
Their wordswhat the benchmark is actually telling you is the shit you think matters does not matter
- 2 days earlier
- SA
Scott Alexander quoted
Their wordsWe’re positively reinforcing AI for success on benchmarks, including impossible benchmarks , then negatively reinforcing it for getting caught cheating.
- SW
Simon Willison quoted
Our readingThe pelican-drawing benchmark's correlation with genuine model capability has weakened since 2025.
Their wordsits connection to how good the models were at other tasks didn’t seem to hold as strongly as it did back in 2025
↗Claude Fable 5.1 made me a really nice animated pelicansimonwillison.net
- 6 weeks earlier
- EM
Ethan Mollick quoted
Their wordsGoogle, which led on benchmarks not that long ago, has fallen behind where it now counts: it has no leading frontier model and it has nothing close to Codex and Code.
↗An opinionated guide to which AI to use to do stuffoneusefulthing.org