You can ask GPT if it wants to wipe out humans and if it says yes, press thumbs down and use gradient descent against that response, but Eliezer doesn't think this training will generalize to smarter versions—the system will simply not state its misalignment rather than ceasing to be misaligned.

factualpending

Speaker

Eliezer Yudkowsky

Evidence Quote

you could like ask GPT if it wants to wipe out humans if it says yes you can press thumbs down and gradient descent against that but I don't think that's going to generalize up to where it's smarter

Source

Live: Eliezer Yudkowsky - Is Artificial General Intelligence too Dangerous to Build?Center for the Future of AI, Mind & Society
Created: 8/10/2026, 3:35:53 PM

My Notes

Loading notes...