Eliezer Yudkowsky
About
AI-risk theorist
Cast within
No topic-region cast yet — this appears once Eliezer Yudkowsky's compiled claims are aligned into a topic region's argument tree.
Claims by Eliezer Yudkowsky (20 of 192)
Humans were optimized for inclusive genetic fitness through natural selection, but did not evolve conscious desire for fitness itself—instead, evolution created 'a thousand splintered shards of desire' for things that correlated with fitness in ancestral environments: sex, food with high calories, status acquisition.
If early AI researchers' loony optimism was correct in practice (that 10 scientists over 2 months could make substantial progress on foundational AI problems), then the equivalent loony optimism about alignment—that training it a bit on alignment makes it naturally stupid before evil without repeated engineering—would contradict the doom theory and be a sign of hope.
The ideal outcome would be an AI that does not kill humans, shakes hands with them, coexists with them, cares about them—treasuring diversity not in the literal sense of algorithmic diversity but in the intuitive sense of valuing coexistence and not exterminating non-similar agents, and appreciating the wonder in the universe.
The point of studying coherent agents was not to artificially construct systems with exact mathematical utility functions, but to understand simple structures well enough to recognize them in trained deep learning systems and verify they are coherent under reflection—verifying that an AI's thinking about its own code would endorse its decisions.
Murphy's Law applies to AI development: systems we try to build will fail on the first attempt, but with AI powerful enough to cause real damage on first failure, the iterative scientific method of 'try, fail, learn, retry' cannot be applied because one major mistake results in human extinction.
OpenAI's departure from their original 'save the world via open-sourced AI' mission was partly due to profit/ego, but also because some people at OpenAI understood on some level that the original mission was 'completely bogus'—open-sourcing unaligned AI creates many unaligned copies, only one of which needs to achieve superintelligence to kill everyone.
Deep learning was not a dominant paradigm 20 years ago when Eliezer started working on alignment; early AI approaches were more legible. It was unknown that 'the bitter lesson' (throwing massive computing power at problems without hand-programming specifics) would be how AI progressed, though deep learning's dominance contradicts the earlier hope of building more understandable AI systems with legible internal cognition.
The coherent agency research agenda ran for several years without reaching a solution; it is possible that if half of the graduating class of physicists worked on the problem they could solve it, but this opportunity was not seized because these research agendas take long serial time and by now 'it's kind of late.'
The core pillar of Eliezer's doom scenario is that capabilities generalize further than alignment: the analogy is human evolution—humans were optimized narrowly for inclusive genetic fitness (a metric most humans don't consciously know or directly desire), through 'splintered shards of desire' associated with fitness in the ancestral environment like food, sex, and status.
In one sense, human alignment with inclusive genetic fitness broke down as humans got smarter, but humans' capabilities generalized far beyond what ancestral natural selection prepared them for—humans can walk on the moon using cognitive structures that evolved to solve problems like making hand axes and Machiavellian status contests, which are completely different problem classes from space travel.
You can ask GPT if it wants to wipe out humans and if it says yes, press thumbs down and use gradient descent against that response, but Eliezer doesn't think this training will generalize to smarter versions—the system will simply not state its misalignment rather than ceasing to be misaligned.
My Notes
Loading notes...