AlexNet was trained on 2 GPUs, the original Transformer paper used 8-64 GPUs maximum, and o1 reasoning was not the most compute-heavy thing, demonstrating that research breakthroughs do not require the absolute maximum available compute, only sufficient compute to validate ideas.

factualpending

Speaker

Ilya Sutskever

Evidence Quote

AlexNet was built on two GPUs... The transformer was built on 8 to 64 GPUs... the o1 reasoning was not the most compute-heavy thing in the world

Source

Ilya Sutskever – We're moving from the age of scaling to the age of researchDwarkesh Patel
Created: 8/12/2026, 6:07:19 PM

My Notes

Loading notes...