Anecdata but, 3 years ago I got an old mining RX 580 4GB for ~120 CAD (about a $100). That card can run almost everything at 1080p and has been used a lot ever since.
Not only were ours undervolted, but also individually tuned for best performance/watt. A very difficult thing to do at my scale since the failure mode is a full machine crash.
Only thing that really degrades is the paste on the heatsink and that's fairly easy to fix.
Next up is just tracking inventory, making changes to the system, etc... this is over 8k individual computers in multiple data centers.
We also added a different class of hardware which was blade based... which increased the individual computers significantly. Ended up with a very cool iPXE boot solution for that.
I also built some pretty cool software to manage it all. It runs on the concept that each machine is an individual worker that knows how to self-heal itself. Even just distributing the software to so many machines reliably, is a challenge.
It has been a fun few years.
Edit: power supplies on the other hand... are a mess. Mostly hand soldered in China... they fail randomly due to the environment they run in. Sometimes, they "die", let rest for a day or two and then fire back up and run just fine.
Why do they reboot? We run on the edge of peak OC tuning performance by default and I've built an automated tuner which downclocks individual cards. This way, they get more stable over time, while maintaining their best possible performance.
Occasionally, we would reset the tunings and then let them auto tune back... this accounted for the seasonal variances because hotter cards are more prone to crashing.
Again, this isn't an actual issue and I have the data to prove it.
You wouldn't want my cards, because they don't have fans. Most people don't have adequate cooling for something like that.
It may still be a lot less 'shock' than normal use, where players have a 15 minute round, then low use for a couple minutes, etc, for hours.. and then turn the card off.
Thermal cycling is known to be bad for electronics-- this is well studied and documented. Sustained high temperatures are also bad, but it's only really bad when the temperatures are really high.
Certainly, thermal cycling can be an issue for electronics in general, but my experience with these specific cards says that it isn't an issue at all. At least certainly not as much as something that should dictate purchasing 'miner' cards or not.
Do you know what causes NVIDIA cards to have their output turn off (black screen) and the fan to go 100%?
Been happening to my 2000-series recently but I don’t know what to try to fix: cooling, PSU, or capacitors…
My guess is a vbios or driver bug. You could also be running into a tuning issue. GPUs are amazingly complex beasts.
Also, 70c is well into the temperature range that will significantly age capacitors.
My point stands: a capacitor that spends most of its life at room temperature except for a few hours at ~60c is going to last significantly longer than a mining card which spends 24x7 at 60-70c, regardless of temperature rating.