Teenagers learn to drive in 10 hours with minimal explicit rewards or verification, instead relying on unsupervised experience, internalized sense of progress, and robust value functions; achieving this in AI would require fundamentally different training approaches than current supervised or RL paradigms.
factualpending
Speaker
Ilya SutskeverEvidence Quote
“It takes much fewer samples. It seems more unsupervised. It seems more robust?... Much more robust. The robustness of people is really staggering.”
Created: 8/12/2026, 6:07:19 PM
My Notes
Loading notes...