We should seek to instill in AI systems robust values like high integrity and honesty, such that when they face requests that seem harmful they refuse to engage; this is analogous to how we try to teach children good values without needing perfect agreement on universal morality.
normativepending
Speaker
Dwarkesh PatelEvidence Quote
“They will refuse to engage in it. Or they'll be honest, things like that.”
Created: 8/11/2026, 6:36:28 AM
My Notes
Loading notes...