The point of studying coherent agents was not to artificially construct systems with exact mathematical utility functions, but to understand simple structures well enough to recognize them in trained deep learning systems and verify they are coherent under reflection—verifying that an AI's thinking about its own code would endorse its decisions.

definitionpending

Speaker

Eliezer Yudkowsky

Evidence Quote

the point of this is not you like artificially construct an agent with that exact utility function...but you could understand the simple structure you were trying to train in and know that it was coherent under reflection

Source

Live: Eliezer Yudkowsky - Is Artificial General Intelligence too Dangerous to Build?Center for the Future of AI, Mind & Society
Created: 8/10/2026, 3:35:53 PM

My Notes

Loading notes...