AlexNet was trained on 2 GPUs, the original Transformer paper used 8-64 GPUs maximum, and o1 reasoning was not the most compute-heavy thing, demonstrating that research breakthroughs do not require the absolute maximum available compute, only sufficient compute to validate ideas.
factualpending
Speaker
Ilya SutskeverEvidence Quote
“AlexNet was built on two GPUs... The transformer was built on 8 to 64 GPUs... the o1 reasoning was not the most compute-heavy thing in the world”
Created: 8/12/2026, 6:07:19 PM
My Notes
Loading notes...