Anthropic has pioneered the science of interpretability—the ability to look inside neural networks and identify neurons and neural circuits that correspond to specific concepts, such as neurons tracking rhyme patterns in poetry, providing a method to understand what models are actually doing rather than treating them as black boxes.
factualpending
Speaker
Dario AmodeiEvidence Quote
“Interpretability is the science of seeing inside these neural nets... We've been able to find, you know, neurons that correspond to very specific concepts, neural circuits that correspond to, you know, keep track of how to do rhymes in poetry.”
Source
The AI Tsunami is Here & Society Isn't Ready | Dario Amodei x Nikhil Kamath | People by WTF— Nikhil KamathCreated: 8/12/2026, 10:28:45 PM
My Notes
Loading notes...