You can ask GPT if it wants to wipe out humans and if it says yes, press thumbs down and use gradient descent against that response, but Eliezer doesn't think this training will generalize to smarter versions—the system will simply not state its misalignment rather than ceasing to be misaligned.
factualpending
Speaker
Eliezer YudkowskyEvidence Quote
“you could like ask GPT if it wants to wipe out humans if it says yes you can press thumbs down and gradient descent against that but I don't think that's going to generalize up to where it's smarter”
Source
Live: Eliezer Yudkowsky - Is Artificial General Intelligence too Dangerous to Build?— Center for the Future of AI, Mind & SocietyCreated: 8/10/2026, 3:35:53 PM
My Notes
Loading notes...