In short
Phil Bell interviews AI safety researcher Lee Sharkey on mechanistic interpretability, the science of reverse-engineering neural networks to understand their internal computation. The core argument: we can build powerful LLMs far more easily than we can explain them, with current understanding perhaps only 1% of what's needed. Borrowing tools from computational neuroscience, researchers use sparse autoencoders to surface interpretable 'features', but still can't trace step-by-step computation.
We can build systems we cannot explain. That gap, between our ability to make powerful AI and our ability to understand it, is the subject of this conversation, and arguably the central problem in AI safety.
Lee Sharkey works on mechanistic interpretability: the attempt to reverse-engineer what’s actually happening inside a neural network, borrowing tools from the field he trained in, computational neuroscience. His sobering estimate is that we currently understand perhaps 1% of how these models work. He walks through how researchers use sparse autoencoders to pull human-readable “features” out of the noise, including Anthropic’s famous Golden Gate Bridge feature, and why that’s progress but nowhere near enough.
If you want to grasp why “we don’t really know how it works” is a literal statement about frontier AI, not a figure of speech, start here. It’s also one of the clearest explanations of interpretability you’ll find for a non-technical audience.
Key takeaways
- Building AI is far easier than understanding it; we may grasp only ~1% of how LLMs work internally.
- Mechanistic interpretability borrows directly from computational neuroscience.
- Models encode concepts as activation patterns across very high-dimensional spaces that defy human intuition.
- Sparse autoencoders find recurring interpretable features (e.g. Anthropic's Golden Gate Bridge feature) but don't explain the full computation.
- Current interpretability is inadequate to verify the safety of superhuman systems; methods must improve and scale.
Listen to the full episode
This is a summary. Listen to the full conversation on your platform of choice, or read the companion essay on Substack.
Frequently asked questions
What is mechanistic interpretability?
How well do we understand how LLMs work?
Why does interpretability matter for AI safety?
People & ideas in this piece
Topics: How AI Actually Works , AI Safety & Alignment