Value functions—intermediate rewards that signal whether a system is on a promising path before reaching a final goal—are fundamental to efficient learning but are underutilized in current LLM training because long-horizon tasks make learning from final outcomes inefficient.
causalpending
Speaker
Ilya SutskeverEvidence Quote
“The value function lets you short-circuit the wait until the very end”
Created: 8/12/2026, 6:07:19 PM
My Notes
Loading notes...