What Eugene Yan thinks about LLMs
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
Everything they publish, on ppll ↗
Eugene Yan did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
4 dated positions, 2025, in their own words. Our reading of what Eugene Yan has said — not written or endorsed by them.
-
Their wordsThe benchmark is human performance, not perfection. We sometimes get requirements for 90%+ accuracy.
↗Product Evals in Three Simple Stepseugeneyan.com 1st of 3 in this piece
-
Our readingThe main advantage of an LLM evaluator over human annotators is scalability, not higher accuracy.
Their wordsIn my opinion, the true benefit isn't higher accuracy than human annotators-it's scalability.
↗Product Evals in Three Simple Stepseugeneyan.com 3rd of 3 in this piece
- 2 months earlier
-
Their wordslanguage models have world knowledge and can eloquently talk about products, but are unaware of our catalog. Also, their recommendations are generic and suffer from popularity bias.
↗Training an LLM-RecSys Hybrid for Steerable Recs with Semantic IDseugeneyan.com 1st of 2 in this piece
- 3 months earlier
-
Our readingLLM-based evaluation methods are more reliable and nuanced than traditional automated metrics
Their wordsThis is why model-based evaluation is increasingly popular-it offers more reliable and nuanced evals than traditional metrics.
↗Evaluating Long-Context Question & Answer Systemseugeneyan.com 2nd of 3 in this piece