In the current generation There are plenty of questions around
- viability of training to inference cascades (the key to extended life) given custom ASICs hitting production like cerebras did early this year.
- energy efficiency of older chips in tight energy environments , just new grid capacity constraints favor running newer efficient chips ignoring perhaps short term(< 1 year) price shock due to war.
- higher MBTF , compared to older GPUs modern nodes are 8 GPU clusters built on 2/3 nm processors depending on HBM memory, the tolerances are much lower especially for training.
- new DCs being spun up are being by up less than ideal conditions due to permitting, part supply and other constraints which will impact operating environment.
Not withstanding, all these issues and even taking a generous 10 year useful life . The expenses dwarf every mega project before it .
Will it be worth the cost of electricity to run them if the flops/watt of newer chips is lower?
If every latest-gen is booked solid and there is still unmet demand, why would you decommission?
A typical node today is 8 GPU node today , you have to keep replacing failed GPUs by cannibalizing parts from other GPUs as nobody is selling new GPUs of that model anymore at higher frequencies.
In addition to outright failure there are higher error rates in computation in graphics it tends to be flickers or screen artifacts and so on.
Azure operated K-80s and P-100s for 9 and 7 years respectively but they were running at 2 GPU nodes and of course were much simpler compared to today’s HBM behomouths on 2/5 nm processor nodes . Google operates their custom ASIC TPUs for about 8-9 years .
With custom inference ASICs like cerebras hitting production the cascading of training NVIDIA chips to inference to get the 5-6 year useful life is also not clear.