If early AI researchers' loony optimism was correct in practice (that 10 scientists over 2 months could make substantial progress on foundational AI problems), then the equivalent loony optimism about alignment—that training it a bit on alignment makes it naturally stupid before evil without repeated engineering—would contradict the doom theory and be a sign of hope.
factualpending
Speaker
Eliezer YudkowskyEvidence Quote
“if that kind of looney-eyed optimism is just correct in practice and you know like just like train it a bit on alignment and it doesn't generalize perfectly to everything but like it gets stupid or much faster than it gets more evil then uh that contradicts the central theory”
Source
Live: Eliezer Yudkowsky - Is Artificial General Intelligence too Dangerous to Build?— Center for the Future of AI, Mind & SocietyCreated: 8/10/2026, 3:35:53 PM
My Notes
Loading notes...