Self-play is most useful for narrow domains like competition and social negotiation, but adversarial setups more broadly (debate, prover-verifier, LLM-as-judge) can create diversity by incentivizing agents to differentiate their approaches from each other.

factualpending

Speaker

Ilya Sutskever

Evidence Quote

self-play... it's only good for negotiation, conflict... debate, prover-verifier... LLM-as-a-Judge which is also incentivized to find mistakes

Source

Ilya Sutskever – We're moving from the age of scaling to the age of researchDwarkesh Patel
Created: 8/12/2026, 6:07:19 PM

My Notes

Loading notes...