korrents
Neel Nanda

Neel Nanda

@neel-nanda · 17 positions · 0 changes of mind

Runs the mechanistic interpretability team at Google DeepMind; worked on interpretability at Anthropic before that.

Everything they publish, on ppll ↗

Neel Nanda did not write this page.

We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.

  1. neural networks

    I'm now fairly pessimistic about ambitious interpretability (i.e. complete reverse-engineering), and I'm excited about model biology (studying qualitative high-level properties of models) and applied interpretability (rigorously doing useful things with interp).

    MATS Applications Open (Due Aug 29)neelnanda.io 1st of 3 in this piece

  2. I'm more agnostic about the best techniques, things like sparse autoencoders are a useful tool, but easy to waste effort using when a simpler method is sufficient or better - start by doing the obvious thing!

    MATS Applications Open (Due Aug 29)neelnanda.io 2nd of 3 in this piece

  3. taste

    My model is that research requires a mix of skills. The day-to-day coding and execution is crucial. But there's also a set of harder-to-learn conceptual skills, collectively called research taste. These skills take a long time to gain because they have poor feedback loops, but they take very little time to use.

    MATS Applications Open (Due Aug 29)neelnanda.io 3rd of 3 in this piece

  4. 3 months earlier
  5. It's a lot easier for someone to engage with an argument if they generated the key steps themselves by answering my questions - if imposed by me, it sparks contrarianism and defensiveness

    Post 51: Socratic Persuasion: Giving Opinionated Yet Truth-Seeking Adviceneelnanda.io 1st of 2 in this piece

  6. No matter how much I know about a domain, the other person will always know far more about their own situation, context, beliefs, skills, preferences, etc than I do.

    Post 51: Socratic Persuasion: Giving Opinionated Yet Truth-Seeking Adviceneelnanda.io 2nd of 2 in this piece

  7. 2 months earlier
  8. AGI

    Having a good research track record is some evidence of good big-picture takes about AGI, but it's weak evidence.

    Post 50: Good Research Takes are Not Sufficient for Good Strategic Takesneelnanda.io 1st of 2 in this piece

  9. Though note that people's opinions are often substantially reflections of the people they speak to most, rather than what's actually true.

    Post 50: Good Research Takes are Not Sufficient for Good Strategic Takesneelnanda.io 2nd of 2 in this piece

  10. 3 years earlier
  11. in general, I expect 90% of a mentoring chat to be kinda useless, and 10% to be particularly high value.

    Post 49: Things That Make Me Enjoy Giving Career Adviceneelnanda.io 1st of 3 in this piece

  12. often the best way to maximise your impact is by pursuing high-risk, high-reward strategies

    Post 49: Things That Make Me Enjoy Giving Career Adviceneelnanda.io 2nd of 3 in this piece

  13. In particular, most EA careers and most careers to do with AI are competitive as fuck.

    Post 49: Things That Make Me Enjoy Giving Career Adviceneelnanda.io 3rd of 3 in this piece

  14. 2 days earlier
  15. In hindsight, the hardest comparisons are in fact the least useful - if I do the fourth most important task before the third most important, that's no big deal, but doing the tenth before the first is a screw-up!

    Post 48: Prioritise Tasks by Rating not Sortingneelnanda.io 1st of 2 in this piece

  16. There's a surprising amount of value in having any order at all (even an arbitrary one like first-come-first-served!) to break past indecisiveness about where to start.

    Post 48: Prioritise Tasks by Rating not Sortingneelnanda.io 2nd of 2 in this piece

  17. 4 months earlier
  18. Someone being smart and competent just means they're right more often, not that they're always right

    Post 47: How I Formed My Own Views About AI Safetyneelnanda.io 1st of 2 in this piece

  19. mathematics

    Mathematicians get feedback re whether there proofs work in a way that, as far as I can tell, moral philosophy doesn't

    Post 47: How I Formed My Own Views About AI Safetyneelnanda.io 2nd of 2 in this piece

  20. 5 days earlier
  21. It's obviously worth it to go on 99 unsuccessful dates if the hundredth results in marriage.

    Post 46: Reward Good Bets That Had Bad Outcomesneelnanda.io 1st of 3 in this piece

  22. Reality is not fully knowable. And thinking harder has costs.

    Post 46: Reward Good Bets That Had Bad Outcomesneelnanda.io 2nd of 3 in this piece

  23. many of the smartest people I know are super insecure and risk averse.

    Post 46: Reward Good Bets That Had Bad Outcomesneelnanda.io 3rd of 3 in this piece