Eugene Yan
Member of technical staff at Anthropic. He has led ML/AI teams at Amazon, Alibaba and Lazada, and writes about LLMs, recommender systems and engineering at eugeneyan.com.
Everything they publish, on ppll ↗
Eugene Yan did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
-
Their wordsThe experiments showed that the system framework matters far more than the underlying model.
- 7 weeks earlier
-
Their wordsnone of this is specific to AI. It's simply how you onboard and work with any new collaborator.
↗How to Work and Compound with AIeugeneyan.com 1st of 3 in this piece
-
Their wordsYou can't delegate what you can't verify, so this requires first defining success criteria and metrics.
↗How to Work and Compound with AIeugeneyan.com 2nd of 3 in this piece
-
Their wordsThe bottleneck has shifted from doing the work to writing clear specs and reviewing outputs fast enough to keep the pipeline moving-the middle is hollowing out.
↗How to Work and Compound with AIeugeneyan.com 3rd of 3 in this piece
- 5 months earlier
-
Their wordsThe benchmark is human performance, not perfection. We sometimes get requirements for 90%+ accuracy.
↗Product Evals in Three Simple Stepseugeneyan.com 1st of 3 in this piece
-
Their wordsAnd human annotators can miss as many as 50% of the defects due to fatigue after looking at hundreds of samples.
↗Product Evals in Three Simple Stepseugeneyan.com 2nd of 3 in this piece
-
Our reading
The main advantage of an LLM evaluator over human annotators is scalability, not higher accuracy.
Their wordsIn my opinion, the true benefit isn't higher accuracy than human annotators-it's scalability.
↗Product Evals in Three Simple Stepseugeneyan.com 3rd of 3 in this piece
- 5 weeks earlier
-
Their wordsBeing right is less than half the battle. You also have to convince others that you're right, and more importantly, convince them to care enough to act on it.
↗Advice for New Principal Tech ICs (i.e., Notes to Myself)eugeneyan.com 1st of 3 in this piece
-
Their wordsTo get to principal, you need to put yourself on the critical path. To be effective as a principal and go beyond it, you need to actively remove yourself from it.
↗Advice for New Principal Tech ICs (i.e., Notes to Myself)eugeneyan.com 2nd of 3 in this piece
-
Their wordsWith great freedom comes great responsibility. You have the autonomy to choose what to work on, but there's the expectation of accountability and impact.
↗Advice for New Principal Tech ICs (i.e., Notes to Myself)eugeneyan.com 3rd of 3 in this piece
- 5 weeks earlier
-
Their wordslanguage models have world knowledge and can eloquently talk about products, but are unaware of our catalog. Also, their recommendations are generic and suffer from popularity bias.
↗Training an LLM-RecSys Hybrid for Steerable Recs with Semantic IDseugeneyan.com 1st of 2 in this piece
-
Their wordsThey excel at predicting what a user will click or buy next, but can't be steered via natural language or reason on their choices.
↗Training an LLM-RecSys Hybrid for Steerable Recs with Semantic IDseugeneyan.com 2nd of 2 in this piece
- 3 months earlier
-
Their wordsPeople find comparing two answers easier than assigning absolute ratings, resulting in greater consistency across annotators.
↗Evaluating Long-Context Question & Answer Systemseugeneyan.com 1st of 3 in this piece
-
Our reading
LLM-based evaluation methods are more reliable and nuanced than traditional automated metrics
Their wordsThis is why model-based evaluation is increasingly popular-it offers more reliable and nuanced evals than traditional metrics.
↗Evaluating Long-Context Question & Answer Systemseugeneyan.com 2nd of 3 in this piece
-
Their wordsSince these datasets are likely already part of model training data, we shouldn't rely solely on them to evaluate our Q&A system.
↗Evaluating Long-Context Question & Answer Systemseugeneyan.com 3rd of 3 in this piece