But I seem to remember they were running 1,000+ 32gb GPUs for 3 months to train it and keeping that infrastructure running day-to-day and tweaking parameters as training continued was the bulk of the 100 pages. It is beyond the reach of anybody but a really big company, at least in the area of very large models, and the large models are where all the recent results are. I wish I was more bullish on algorithm improvements meaning you can get better results on less hardware; there will definitely be some algorithm improvements, but I think we might really need more powerful hardware too. Or pooled resources. Something. These models are huge.