MuZero models for Atari games could not be used to train a single policy to play both chess and Go and other games; instead, each game had to be trained in a specialized way, suggesting information constraints may limit generalization.

factualpending

Speaker

Dwarkesh Patel

Evidence Quote

You couldn't, using that framework, train a policy to play both chess and Go and some other game. You had to train each one in a specialized way.

Source

Richard Sutton – Father of RL thinks LLMs are a dead endDwarkesh Patel
Created: 8/11/2026, 6:36:28 AM

My Notes

Loading notes...