I worked at a few data centers on and off in my career. I got lots of hardware for free or on the cheap simply because the hardware was considered “EOL” after about 3 years, often when support contracts with the vendor ends.
There are a few things to consider.
Hardware that ages produce more errors, and those errors cost, one way or another.
Rack space is limited. A perfectly fine machine that consumes 2x the power for half the output cost. It’s cheaper to upgrade a perfectly fine working system simply because it performs better per watt in the same space.
Lastly. There are tax implications in buying new hardware that can often favor replacement.
But no, there’s none to be found, it is a 4 year, two generations old machine at this point and you can’t buy one used at a rate cheaper than new.
These things are like cars, they don't last forever and break down with usage. Yes, they can last 7 years in your home computer when you run it 1% of the time. They won't last that long in a data center where they are running 90% of the time.
Even assuming your compute demands stay fixed, its possible that a future generation of accelerator will be sufficiently more power/cooling efficient for your workload that it is a positive return on investment to upgrade, more so when you take into account you can start depreciating them again.
If your compute demands aren't fixed you have to work around limited floor space/electricity/cooling capacity/network capacity/backup generators/etc and so moving to the next generation is required to meet demand without extremely expensive (and often slow) infrastructure projects.
Two years later, H900 is released for a similar price but it performs twice as many TFlOps/Watt. Now any datacenter using H900 can offer the same performance as NeoCloud Inc at $5/month, taking all their customers.
[all costs reduced to $/minute to make a point]
Current estimates are about 1.5-2 years, which not-so-suspiciously coincides with your toy example.
For servers I've seen where the slightly used equipment is sold in bulk to a bidder and they may have a single large client buy all of it.
Then around the time the second cycle comes around it's split up in lots and a bunch ends up at places like ebay
A lot of demand out there for sure.
Yes. I'd expect 4 year old hardware used constantly in a datacenter to cost less than when it was new!
(And just in case you did not look carefully, most of the ebay listings are scams. The actual product pictured in those are A100 workstation GPUs.)
Rack space and power (and cooling) in the datacenter drives what hardware stays in the datacenter
I have not seen hard data, so this could be an oft-repeated, but false fact.
If this was anywhere close to a common failure mode, I'm pretty sure we'd know that already given how crypto mining GPUs were usually ran to the max in makeshift settings with woefully inadequate cooling and environmental control. The overwhelming anecdotal evidence from people who have bought them is that even a "worn" crypto GPU is absolutely fine.
Another commonly forgotten issue is that many electrical components are rated by hours of operation. And cheaper boards tend to have components with smaller tolerances. And that rated time is actually a graph, where hour decrease with higher temperature. There were instances of batches of cards failing due to failing MOSFETs for example.
Even in amateur setups the amount of power used is a huge factor (because of the huge draw from the cards themselves and AC units to cool the room) so minimising heat is key.
From what I remember most cards (even CPUs as well) hit peak efficiency when undervolted and hitting somewhere around 70-80% max load (this also depends on cooling setup). First thing to wear out would probably be the fan / cooler itself (repasting occasionally would of course help with this as thermal paste dries out with both time and heat)
Not sure I understand the police raid mentality - why are the police raiding amateur crypto mining setups ?
I can totally see cards used by casual amateurs being very worn / used though - especially your example of single mobo miners who were likely also using the card for gaming and other tasks.
I would imagine that anyone purposely running hardware into the ground would be running cheaper / more efficient ASICS vs expensive Nvidia GPUs since they are much easier and cheaper to replace. I would still be surprised however if most were not proritising temps and cooling
It's like if your taxi company bought taxis that were more fuel efficient every year.
You kind of have to.
Replacing cars every 3 years vs a couple % in efficiency is not an obvious trade off. Especially if you can do it in 5 years instead of 3.
It can make sense at a certain scale, but it’s a non trivial amount of cost and effort for potentially marginal returns.
I’m just pointing out changing it out at 5 years is likely cheaper than at 3 years.
Company A has taxis that are 5 percent less efficient and for the reasons you stated doesn't want to upgrade.
Company B just bought new taxis, and they are undercutting company A by 5 percent while paying their drivers the same.
Company A is no longer competitive.
The scenario doesn't add up.
If company A still has debt from that, company B has that much debt plus more debt from buying a new set of taxis.
Refreshing your equipment more often means that you're spending more per year on equipment. If you do it too often, then even if the new equipment is better you lose money overall.
If company B wants to undercut company A, their advantage from better equipment has to overcome the cost of switching.
They both refresh their equipment at the same rate.
I wish you'd said that upfront. Especially because the comment you replied to was talking about replacing at different rates.
So your version, if company A and B are refreshing at the same rate, then that means six months before B's refresh company A had the newer taxis. You implied they were charging similar amounts at that point, so company A was making bigger profits, and had been making bigger profits for a significant time. So when company B is able to cut prices 5%, company A can survive just fine. They don't need to rush into a premature upgrade that costs a ton of money, they can upgrade on their normal schedule.
TL;DR: six months ago company B was "no longer competitive" and they survived. The companies are taking turns having the best tech. It's fine.
Isn't that precisely how leasing works? Also, don't companies prefer not to own hardware for tax purposes? I've worked for several places where they leased compute equipment with upgrades coming at the end of each lease.
who cares? that's the beauty of the lease. once it's over, the old and busted gets replaced with new and shiny. what the leasing company does is up to them. it becomes one of those YP not an MP situations with deprecated equipment.
That's where the analogy breaks. There are massive efficiency gains from new process nodes, which new GPUs use. Efficiency improvements for cars are glacial, aside from "breakthroughs" like hybrid/EV cars.
It's not like the CUDA advantage is going anywhere overnight, either.
Also, if Nvidia invests in its users and in the infrastructure layouts, it gets to see upside no matter what happens.
(1) We simply don't know what the useful life is going to be because of how new the advancements of AI focused GPUs used for training and inference.
(2) Warranties and service. Most enterprise hardware has service contracts tied to purchases. I haven't seen anything publicly disclosed about what these contracts look like, but the speculation is that they are much more aggressive (3 years or less) than typical enterprise hardware contracts (Dell, HP, etc.). If it gets past those contracts the extended support contracts can typically get really pricey.
(3) Power efficiency. If new GPUs are more power efficient this could be huge savings on energy that could necessitate upgrades.
Companies can’t buy new Nvidia GPUs because their older Nvidia GPUs are obsolete. However, the old GPUs are only obsolete if companies buy the new Nvidia GPUs.
This doesn't mean much for inference, but for training, it is going to be huge.