What to do with all that computing power now. I guess AI.
What to do with all that computing power now. I guess AI.
If higher or close to retail - why buy used gpu ?
Let miners literally drawn in debt.
Thermal cycling common to gaming uses causes mechanical failures in chips, between chips and the board, and the board - but electrolytic capacitor lifetime drops dramatically as temperature goes up.
Not much thermal cycling in desktop systems seeing mostly "productivity" use that have proper ventilation.
memory has to stay as cool as possible or else the overclock will be unstable.
The (large) GPU mining operations run at rather high ambient.
That holds (? used to hold?) true from anything to GPUs to individual capacitors, the repetition of the heat cool cycles is the primary cause of wear.
I'm not sure whether constant running is worse for wear and tear vs power cycling but power cycling is typically what precedes a failure.
This isn't true at all.
That failure mode is more, it works fine when initially turned on, then starts failing/flickering on and off once it has gotten to temperature.
Also in this case we have mechanical components under constant usage (fans), and the effects of long term heat on the thermal compound between the die and cooler plates to consider.
Ethereum mining uses a lot of vram, perhaps the overuses of the memory damages it?
There are probably about a dozen people on the entire planet who actually understand how and why GPUs fail, and tens of millions who think that the fact they can build a PC qualifies them to present a regurgitated theory from Reddit as "expert opinion".
This seems like an underestimation. I would bet there are mid-level hundreds able to articulate at a precise level at least, and dozens of thousands more who are sufficiently familiar with the design and shortcomings modern GPUs and chips who would be able to give an approximately decent answer from first principles.
That said, despite my guess being four orders of magnitude larger than yours, the odds of actually encountering such a person among the 7 billion available is practically the same.
I know people who are one-of-a-few experts in technology that is critical to a particular industry, and there's not always good succession plans in place for when they retire.
I could absolutely be wrong in this case - it might be well enough understood that you don't need decades of raw experience to begin to scratch the surface - but at the same time there really are so many advanced technologies held together by one or a few experts.
I guess it’s a power law distribution like so many other things.
Appears to me the effects would be less than the pandemic or the Russian war.
Also having a cheap phone from 2015 is quite an outlier. I absolutely get wanting to moderate tech use and I make full use of the digital wellness limiters to cut down on social media and interruptions, but phones are just so damn useful for just about everything in modern life. Throwing the baby out with the bathwater a bit there, no?
Honestly, I'm not sure how true that is. I got rid of my smartphone just before my daughter was born (she's turning 2 this month) and there has never been a single point where I've regretted it.
That's not to say I avoid Android entirely. I've still got a tablet at home for Netflix/Chromecast, and we'll use my wife's phone in the car for Google Maps, but that's about it.
There really is no need to be 100% connected 24/7.
You are basically reliant on your wife for accessing any kind of digital service without a website, communicating with anyone, or any computing whatsoever when not at home. For most that is giving up on staggering utility, but I know people who prefer to ride horses still, so different strokes of course.
I still use computers and the internet, including Android devices, at home and work (quite non-trivially, actually - I am my family's sole provider and 100% of that income comes from a mix of online business and remote consulting) and I'll still talk to friends, family and employees via SMS or occasionally voice.
Genuinely, man, as someone who has actually done it, you miss out on absolutely nothing and gain so much.
Nobody would deny that, at least not in Europe, North America etc. But I don't think I miss a lot of "modern civilization" by not participating in evey aspect of surveillance capitalism and used mostly apps from F-Droid.
Remember that it was possible to fly to the moon in 1970 with less than a Megabyte of memory. The efficiency of additional resources poured into scientific seems to go down all the time. Not at all sure it is worth ruining the planet by ever increasing resource usage for small gains in knowledge.
- Weather prediction.
- Other aerodynamic and hydrodynamic modeling, used in a lot of industry.
- Image recognition.
- Voice recognition / captioning.
- Machine translation.
- I suppose also stuff like large-scale mechanical modeling, protein folding models, etc.
That is absolutely bigger than the pandemic or the Russian invasion.
However, what probably wouldn't work as well without GPUs is video production and consumption. Linux users knew that for years, when GPUs were poorly supported. Doing on the CPU does not give you the same resolution/framerate and requires more energy. Some might claim that watching less Youtube/Tiktok and reading more books or exercise outside instead would be good for "modern civilization".
Certainly there are many more nuances to wear patterns than that, but that was kind of the basic question I was asking. But as I said originally I was asking from a position of interested though ignorant curiousity, so my assumption that most hardware designers would need to understand this sort is stress might be... discontinuous with the factual nature of the physical & human experiential phenomena relevant to the question. (e.g., stupid assumptions) :)
I am not saying this makes them instant experts on GPU failures. On the other hand I also don't believe GPU failures are that special.
https://semiaccurate.com/2009/08/21/nvidia-finally-understan...
>On July 2, 2009, the date being ironically a year after the notorious 8-K that publicly kicked off bumpgate, the company put up a job listing for a “DIRECTOR OF PACKAGE TECHNOLOGY”.
If you can't accelerate the failure mechanism in a well defined way, you cannot lab test long-term reliability other than just letting the system run. So if you want to guarantee 10 yrs of lifetime, you need to run it for 10 yrs. Obviously, that doesn't work for most products. Instead you have to rely on models, which, for a new node, might not be proven very well.
Bottom line was this vendor shipped products that nobody could really be sure would last the promised lifetime. This was regarded as a top silicon fab in the world. Back then, finer geometries were much harder because of new failure mechanisms. I don't know state of art today, but I would not be surprised if measuring reliability is still not a really hard problem for new nodes.
Edit: I am just referring to silicon issues, the rest of the PCB has its own set of issues totally separate. Generally however, those are better understood than a new process node.
And even if they did, the capacitors in the VRMs of these cards are full of electrolytic capacitors which do degrade over time due to heat and chemical reactions occurring between the plates and the electrolyte as well as the evaporation of electrolyte, they not only age, but age faster over time as the degraded capacitors are put under increased stress.
Besides, the assumption of running these cards underclocked and undervolted might not be true - if the folks running those obtained the electricity illegally, or incredibly cheaply, it might make more sense to run these cards hotter.
Everything considered, these second hand GPUs are not even that cheap, I saw 5700 XTs going for $250 on ebay, with a comparable brand new 6600 costing about $300.
Edit: The fans also have suffered serious wear and tear and are likely to fail. Considering these are often custom, you might not even be able to replace them.
Underclock/volting cards to get a consistent higher net yield is standard practice in crypto mining, just as ovectclicking/volting is standard practice in enthusiast gamer circles. The latter being way more damaging for the GPU. Those that obtain the electricity illegally, or incredibly cheaply would not be the ones selling cards now.
As for fans, the things that reduce the fan's lifetime the most is dust buildup and high positive pressure configurations. Both of these are typical for bedroom gamer setups. Miner cards run in cleaner environments and open frames.
I would definetly buy a GPU from a miner over one from a gamer.
> Thermal cycling such as is typical in a gaming PC is far more damaging than constant temperatures.
the reasoning is probably pretty obvious, but what i heard is the constant expansion and contraction from heating and cooling the chips and connections is what causes the damage im those casessome gpus are rather hard to get a replacement fan for because it's not like they use an off the shelf modular 60 or 80mm 12vdc fan.
The higher end card companies in some cases will just ship you a replacement fan if your email to them is pleasant. Or sell it to you for a modest price.
I just looked and I can get replacement fans for my specific GPU for less than $20 including shipping.
The bigger problem is VRM wear, you never know when it will fail. Said TIM drying out affects the VRM components the most. One failed capacitor or resistor and the whole GPU can get fried (seen it happen many times).
Is there a good way to spot this sort of thing from a distance when evaluating stuff off eg eBay?
SMD elements can fail either open or closed/shorted, the latter is more likely to cause a worse failure but it's not exclusive. Due to protection circuits elsewhere they can fail without visible marks so it's hard to diagnose.
Thermal pads dry out because of heat first, their quality second, with load coming in last imo. So the fact that it's been used for mining alone doesn't mean much.
The conditions under which a card was used and the time matter more - if only sellers told you the average temperature, case type, ambient temperatures, whether the cards were undervolted/overclocked and how long they actually ran...
As far as I can see, the answer is yes[1]:
The relationship between integrated circuit failure rates and time and temperature is a well established fact. The occurrence of these failures is a function which can be represented by the Arrhenius Model. Well validated and predominantly used for accelerated life testing of integrated circuits, the Arrhenius Model assumes the degradation of a performance parameter is linear with time and that MTBF is a function of temperature stress.
However, the dramatic acceleration effect of junction temperature (chip temperature) on failure rate is illustrated in a plot of the above equation for three different activation energies in Figure 2. This graph clearly demonstrates the importance of the relationship of junction temperature to device failure rate. For example, using the 0.99 ev line, a 30° rise in junction temperature, say from 130°C to 160°C, results in a 10 to 1 increase in failure rate.
Wikipedia has an overview of the the Arrhenius equation[2].
Now, as you probably know most complex chips like GPUs and CPUs have built-in thermal management which prevents the junction temperature to rise above some limit, which one would assume is set at a point which reasonably guarantees a decent lifetime.
However, according to this, chips that have experienced less heat should, on average, live longer.
* Dried up electrolytic capacitors
* Bad solder joints exacerbated by lead free solder.
* Failed high current transistors (typically regulation mosfets or triacs in appliances)
Failure of low current handling chips is pretty rare.Gamers are the ones destroying GPU's, not professional miners who tune their systems properly.
DRAM is pretty durable but the controller, bus pins, and other parts would be the wear/tear in my opinion. Those other bits inside the DRAM could very well wear out.
Debatable.
A lot of these miners are keeping their GPU-stations inside of a garage with no air-conditioning (so 90F+ temperatures, maybe 100F+ ambient). These miners are going for the cheapest-of-cheap conditions, they weren't running a datacenter, but a mining operation.
I'm willing to bet that the vast majority of them never paid attention to airflow issues. Closets, garages, maybe even sheds.
While GPU cores are underclocked and undervolted (but not necessarily), the RAM is going to be processing at full tilt as quickly as possible, possibly overclocked and overvolted and therefore at even higher temperatures.
--------
Finally, buying used means having a very ambiguous warranty. You might not have anyone to send the GPUs to. Its generally not a good idea, especially now that new-prices of GPUs is below MSRP.
Honestly, we're not getting a temperature graph of these GPU's life. We're only getting assurances from some cryptobro that they treated their GPUs correctly (and no offense to the cryptobro community, but... there's a lot of lying scammers out there). Its not a trustworthy community at all.
The only place this doesn't apply is gaming, where you're trying to get realtime interactive results while economizing on both cost and energy. But for every other use of a GPU, you can't go wrong with "throwing more compute at the problem."
I'd buy a used one if the price were commensurate with the risk. Say 20% retail.
This
https://www.youtube.com/watch?v=1T0npiqjEWQ
Apparently GDDR6X and HBM memory might degrade from Ethereum mining
This is why it's important to keep battering the crypto advocates any time they promote anything linked to proof-of-waste. Cryptocurrency is the "paperclip maximiser" that will consume scarce energy and hardware manufacturing resources, and turn them into "grey goo" which people have been conned into believing is worth something.
For AI, the lack of memory is pretty limiting.
Having a ton of 8 GB or 10 GB memory GPUs (3070/3080) is annoying and often just not even possible to use effectively without designing specifically to it. If they get really really cheap maybe they'll be some tools or models that are built to leverage scenarios like that, it's not impossible, but for now it's probably just the cheap 3090s (24GB of memory) that will benefit AI.
A couple weeks of training and I break even.
It perhaps isn't the end of the world if your gaming session errors out due to the service being based on used crypto mining GPU cards - especially as the service subscription cost approaches $0/month.
Margins on pre-builds have been shrinking to almost nothing, I build my own PCs because I see it as a hobby, but I could hardly recommend it nowadays to someone who's looking for a machine, not a new hobby.
You say it like people didn't buy GPUs before crypto. What do you think happened to all those GPUs AMD and nVidia have made every year for the past decade or so?
CEO of Kryptex here. We have an answer:
Expect price to dive deeper lol