This points to an interesting future for foundation models. This is an 18x cost reduction in only 2 years. Either foundation models are going to get much bigger, or variations will become common.
An eight GPU DGX-1 server cost ~149k$ back then (googled news postings). A current gen DGX H100 is 520k$ with 5 years of support. Of course it holds 5x the memory, plus GPUs and interconnect are much faster. But when comparing costs, take price hikes into account.
For writing code you don't care about feeding world history to your model. So a smaller model might be better at a specialized task
Sure, having a big multi-modal-model is great, but by having specialized models you can spread tasks better