Anthropic has pioneered the science of interpretability—the ability to look inside neural networks and identify neurons and neural circuits that correspond to specific concepts, such as neurons tracking rhyme patterns in poetry, providing a method to understand what models are actually doing rather than treating them as black boxes.

factualpending

Speaker

Dario Amodei

Evidence Quote

Interpretability is the science of seeing inside these neural nets... We've been able to find, you know, neurons that correspond to very specific concepts, neural circuits that correspond to, you know, keep track of how to do rhymes in poetry.

Source

The AI Tsunami is Here & Society Isn't Ready | Dario Amodei x Nikhil Kamath | People by WTFNikhil Kamath
Created: 8/12/2026, 10:28:45 PM

My Notes

Loading notes...