The point of studying coherent agents was not to artificially construct systems with exact mathematical utility functions, but to understand simple structures well enough to recognize them in trained deep learning systems and verify they are coherent under reflection—verifying that an AI's thinking about its own code would endorse its decisions.
definitionpending
Speaker
Eliezer YudkowskyEvidence Quote
“the point of this is not you like artificially construct an agent with that exact utility function...but you could understand the simple structure you were trying to train in and know that it was coherent under reflection”
Source
Live: Eliezer Yudkowsky - Is Artificial General Intelligence too Dangerous to Build?— Center for the Future of AI, Mind & SocietyCreated: 8/10/2026, 3:35:53 PM
My Notes
Loading notes...