My deep learning rig (2022)
nonint.com
nonint.com
I bought 21 R12 Aurora Alienwares and turned a side-office in my garage into a hot-ass crypto farm with two swamp coolers, four 20A circuits, a 15A circuit, and a bunch of surge protectors. I was always afraid to move something in fear that I'd overload a circuit due to the effort in calculating which power supply was powering which machine(s).
I gave away the 3080s to friends and family and kept the 6 3090s.
Will you ever see positive return on investment from building this or is it feeding the nerd chute?
>> crypto
Money in the stock market (or crypto market but similar) is probably the only way to justify tens of thousands of dollars worth of GPUs that become obsolete quickly.
The other thing I can think of is Reseach scientists who use it for research. But funded research scientists that I know generally don't get hardware to take home. They usually either get tons of cloud credits or time on a supercomputer to do their stuff.
I sold a 4yo gpu for more than I bought it for not long ago.
It's the only thing in a pc that doesn't depreciate much these days, besides maybe the case.
What's a "QS" model?
Okay, I searched and it means "Qualification Sample", i.e. a grey market, non-production grade CPU.
Never encountered this initialism before, despite being a CPU aficionado. Hope this saves you some frustration!
10gbe is very cheap now, but I guess that's not enough?
BTW: echo to the author, PSU and in the U.S. (120v) is a major issue why I am limiting to 4-GPUs. Also, it seems 3090 still have NVLink support, wondering why the author haven't put that up. From what I experienced, NVLink does help if you run data parallel training.
Edit: and you'd have to factor in that cooling power as well into the running costs.
I'd honestly be a bit suspect.
NVLink is about ~$100, so for these cases, I guess you shouldn't expect it to be "more than buying another card" type of improvements.
Once you have larger batch size and gradient accumulation, DDP won't be improved by NVLink I believe (the all-reduce traffic on gradients will be small comparing to your computation overhead).
Enthusiast / small business / entry level enterprise gear will get you there, but you’re looking at several hundred dollars per port.
200GbE and 400GbE is still totally unaffordable for anything remotely personal, IMO.
When looking at cloud GPU availability and current trends (no one except some enthusiasts and bigtech is finetuning and serving on a large scale yet and results keep getting better and better), I fear we will run into a situation where GPUs will be extremely expensive and hard to come by until supply catches up?
I ordered a high end PC with 4090 for the first time in years (normally would always prefer cloud even if more expensive) because I want to be on the safe side. What do you think, is this irrational and just a bubble thing?
The real question is whether NVIDIA will maintain its lock on the market, or if vendor-agnostic Torch will help commoditize the segment.
Is this a competitor to CUDA?
It's pretty astonishing to me AMD has neither a proper math library (like MKL) nor gpu compute library (like cuda).
My fear is that Nvidia trying to anchor GPU prices to crypto levels whether there is TSMC shortage is not.
If AMD were better supported, it would be most economical to use 4x MI60s for 128GB using an Infinity Fabric bridge. However, in order to get to the end of such a journey, you would have to know something.
This would severely limit training using model parallelism.
For data parallel where the full model fits on each card and the batch size is just increased it wouldn't matter as much, and maybe that is the primary use for this.
I wonder how this is dealt with on vast.ai rentals. Because there is a huge difference if I needed 7x 3090's where I need all 168GB to load weights on a single giant LLM model vs. just wanting to run 4GB Stable Diffusion in parallel inference with a massive batch size....
See here [0] about just the difference between having NV-link between cards or not, and the 23% increase in training speed, and the note that the peak bandwidth between 2x 3090s with the link is a peak of 112.5 GB/sec.
[0]: https://huggingface.co/docs/transformers/v4.31.0/en/perf_har...
Now look at PCIe 3.0 speeds, which would be what any two cards talking to each other would need to use thru your risers- only 15.754 GB/s on x16 and only 7.877 GB/s if you are on a x8 riser.
For some non-ML things that I use GPU's for (CFD), the interconnect / memory access bandwidth is the bottleneck, and the simulation time literally scales near linearly with the PCIe bandwidth I have between cpu lanes and the cards.
They also have a little stat that lists 'per-GPU' bandwidth in GB/s and what PCIe version and speed is being used. So they must run some tests beforehand to gauge this. When I look on there now it varys setup to setup. I see people running quad 4090's on PCIe4.0 x16 with 24GB/s bandwidth between them, some people running x8, some on PCIe3.0 with 11 GB/s, even someone with quad 3090's but all on PCIe 2.0 x1 slots with the bandwidth reading 0.3GB/s!!! (likely an old mining rig with those x1 slots)
The have a overall "DLperf" deep learning performance score which is an performance metric and it seems like if you want to get users you set your price per hour at something competitive in line with the market given that metric.
For instance, the guy with the quad 3090s on x1 slots actually has an hourly price set HIGHER than everyone else... even with horrible DLPerf score. No one is using that, ever.
Doubly so for Intel. They literally have no pro market to lose with a 32GB A770, and everything to gain from momentum for their stack.
The W7900 is officially supported by ROCm on Windows. On Linux, the W7900 is enabled, though not officially supported.
Thats not A/H100 terrible, but its still out of reach for many nonprofessional ML users/tinkerers.
Nvidia is Nvidia as needs to preserve their pro tier stratification. But AMD has less to lose, and Intel has nothing to lose by pricing them more like gaming cards.
Though looking around you can get an MI60 on ebay for $500, seems really good for 32gb and still seems to be supported by rocm (as it's just an MI50 with more hbm). Looks to be the cheapest way of getting a GPU with that region of memory support, though things like FP16 speed and BF support suffers compared to later generations. Though from what I've seen most "home" ML tasks are often memory limited pretty hard before ALU limitations kick in. And no idea if a home user would be able to use the IF links either for helping multi-gpu either.
And that includes energy costs so I assume the OP has a cheap source of power. Here in NL I could not do this profitably, even off solar power it would be more efficient to sell that power to the grid than to use it to drive a GPU rig.
With the proper plumbing, you could hook up your water heater to it as well.
Co-generation is in fact a pretty good idea when you start running large computers at home. But in summer...
Guilty as charged.
That was never the case with a crypto currency.
That's a flat out lie and you know it.