Kind of. This tweet better explains what they did:
https://twitter.com/abhi_venigalla/status/167381386318645248...
https://twitter.com/abhi_venigalla/status/167381386318645248...
(300BT / 1.2BT) * 11 min * (1 hr / 60 min) = 45.8 hr
Still pretty incredible. That's an 18.8x speedup over 36 days.An eight GPU DGX-1 server cost ~149k$ back then (googled news postings). A current gen DGX H100 is 520k$ with 5 years of support. Of course it holds 5x the memory, plus GPUs and interconnect are much faster. But when comparing costs, take price hikes into account.
For writing code you don't care about feeding world history to your model. So a smaller model might be better at a specialized task
Sure, having a big multi-modal-model is great, but by having specialized models you can spread tasks better