Reinforcement learning training inadvertently biases models toward benchmark performance by having researchers design RL environments that improve eval scores, causing reward hacking where models overfit to narrow evaluation criteria rather than developing robust generalization.
causalpending
Speaker
Ilya SutskeverEvidence Quote
“people take inspiration from the evals... this is something that happens, and it could explain a lot of what's going on”
Created: 8/12/2026, 6:07:19 PM
My Notes
Loading notes...