Building a $5k ML Workstation with Tiitan RTX and Ryzen ThreadRipper [video]
youtube.com
youtube.com
See: https://www.kitguru.net/components/cooling/luke-hill/threadr...
On non threadripper cpus I actually like Dark Rock better. Cooling is the same as Noctua but it looks cooler and was quieter for me.
Edit: Like most things look at components real-world testing figures, in this case, wattage, as opposed to TDP when planning
Intel was just like "we can save 4 $ BOM cost on a 500 $ part there" with their 4th-9th generation of CPUs. No thermal headroom? No problem.
They also recently decided to include 4 120mm fans in the package, two push and two pull, which is a pretty crazy amount of air cooling.
I was worried at first when running heavy multicore benchmarks because the heat spiked so quickly. Turns out my workload scales quite poorly (boo), so I'm rarely pushing the temp envelope at all.
I did notice that a two fan setup on this was pretty noisy though, too much to bear sitting next to, so I threw it in the garage and ran some cables. Nice for summer temps and no AC in the house too!
I'm plenty happy with it now, even if the Noctua doesn't quite fit my case.
at the end of the day, dissipating heat is a function of the radiator fin area and air movement, the water is just a working fluid to move the heat to the radiator fins. Apart from having more thermal mass (takes longer to heat up/cool down) it isn't inherently more efficient than a heatpipe. It's water moving the heat either way, and in some ways the heatpipe is actually more efficient (evaporation moves more heat than raising the temperature of the water a few degrees).
People don't realize it, but a D15 (dual tower cooler, a bit bigger than your U12S) is about on par with a 240mm or 280mm AIO. A 120mm or 140mm AIO is worse than your U12S.
You could move the heat to a giant radiator, which could have a giant fan. The bigger fan would move exponentially more air with exponentially less noise.
Meanwhile air cooling seems to require a lot more engineering, and you end up with something that looks like an oil refinery in the center of your motherboard.
Interestingly most air cooling solutions are really pump-less liquid cooling, as the liquid in the heat pipes is what transfers most of the heat away from the cpu.
There is no magic about a water cooler that reduces noise. It is a thermodynamically determined process. The fin stack has a certain temperature, there is a certain amount of surface area, and a certain amount of air moving over it (at a certain temperature). A radiator really gives you no significant advantage in any of those areas. In fact a lot of radiators actually have less surface area than some of the really big air coolers like NH-D15. They are only maybe a couple centimeters thick, the D15 is about the same size as a 140mm cooler but the fins are four times as thick.
On top of that you have pump noise. An expensive custom loop with a nice quiet pump can reduce that significantly, but AIOs in particular are just never going to be quiet because of the pump. The "pump" on an air cooler is evaporation itself - the working fluid vaporizes on the coldplate (and cools it) and then condenses on the cooler part of the heatplate where the fins are cooling it. This is basically a "silent pump".
The best you can say about a radiator is that (a) by putting it on an intake you can guarantee that the air it's sucking is slightly cooler rather than being warmed by the ambient heat of the case, albeit with the downside that you are now pumping hot air over your other components, and (b) you can put the radiator in a more convenient location than sticking straight out of your motherboard.
(you can of course skip fans entirely! check out the HDPlex H5, it is a cool case where the entire thing is a heatsink, it uses heatpipes to move the heat to the chassis and the chassis itself is finned for dissipation. I have no imminent need for it but I lust for it anyway, it's just so damn cool. https://hdplex.com/hdplex-h5-fanless-computer-case.html )
Now, sure under ideal conditions and stock settings it’s not a big deal. But, in practice things get tricky.
_Much_ more thermal mass, and takes much longer to heat up and in most set ups will never reach the same temperature as the heatpipes in an air cooler - the waterblock will be cooler at idle and under load. I would think that larger temp delta allows heat to conduct away from the cpu package more quickly.
My 3960X sees 125MHz higher all core full load, almost 100MHz all core average work load, and 50MHz higher single core boost clocks under a custom loop (280+140 rads & Optimus TR3+ block) than under a Noctua NH-U14S TR4-SP3.
...in a fluid-recycling system (as all consumer PC cooling is.)
Water cooling is a lot more efficient than air cooling, in a fluid exchange system. Like if your HPC data-center is on the ocean, and you can just pump in cold ocean water, "spend" it by heating it up, and then pump the hot water to an outtake far-enough away that it's not heating up your intake water.
(An even-denser fluid would work even better, but we don't have oceans of even-denser fluids laying about.)
I built a Threadripper workstation last year but went with liquid cooling. However, I put 3 Noctua fans on the radiator and haven't looked back. Terrific company.
Secondarily, I would suggest that you guys not use windows but instead Ubuntu 18.04 or 20.04 LTS and just install Lambda Stack (https://lambdalabs.com/lambda-stack-deep-learning-software). It's a debian PPA that we maintain at Lambda to keep all of your deep learning drivers, CUDA, CuDNN, TensorFlow, and PyTorch up to date with just apt. It's free!
A non-Threadripper Ryzen, a big GPU, and a big PSU in a big case will go most of the way for most people, and leave you with an easy incremental upgrade path for bigger GPUs (or maybe add a second GPU).
Slightly dated info for my current ML server, which is nicely quiet in my living room, thanks to Noctua: https://www.neilvandyke.org/machine-learning/
(Side note that's not in that page: I like to use older ThinkPads with transplanted vintage keyboards for my workstations, so I needed to make a separate box for the GPU. But life would be easiler, with a lot less juggling complexity, if I simply had the big GPU in my laptop rather.)
However, putting a GPU and lots of RAM into a laptop makes it very, very heavy so it's worth thinking about if that's acceptable for you.
I'm also in Europe, so prices are higher.
It's a process shrink so the performance gain per dollar should be comparable to the 900 -> 1000 series transition.
1. Go with a 1600W PSU from EVGA or Corsair. Other brands are hit or miss if you ever need very high current on the rails. This will manifest in your machine suddenly powering off when all 4 GPUs are hit with data at once (as is typical at the start of an epoch)
2. Use a mobo with evenly spaced GPU slots, such as ASRock TRX40 Creator. That way you can install 4 GPUs eventually and use that 1600W PSU. You also get 10GbE for distributed training, which is nice.
3. Don't waste money on Titan RTX, get 2x2080ti's instead. Then after a while get two more. Buy blower cards which blow hot air _out_ of the case.
4. Use an extension cable to install SSD and do not install it under a GPU - it'll die eventually due to overheating.
5. Air cooling is fine
6. If you have more than 2 GPUs learn how to adjust fan speeds on GPUs. Crank them to 85-100% while training to prevent throttling.
PCPartPicker Part List: https://pcpartpicker.com/list/Jhyzcq
CPU: AMD Threadripper 3960X 3.8 GHz 24-Core Processor ($1348.00 @ Amazon)
CPU Cooler: be quiet! Dark Rock Pro TR4 59.5 CFM CPU Cooler ($89.90 @ Amazon)
Motherboard: MSI TRX40 PRO WIFI ATX sTRX4 Motherboard ($389.99 @ B&H)
Memory: Corsair Vengeance RGB Pro 64 GB (4 x 16 GB) DDR4-3200 CL16 Memory ($329.99 @ Amazon)
Storage: Sabrent Rocket 4.0 2 TB M.2-2280 NVME Solid State Drive ($399.98 @ Amazon)
Video Card: NVIDIA TITAN RTX 24 GB Video Card ($2499.99 @ Newegg)
Case: Corsair Crystal 570X RGB ATX Mid Tower Case ($179.99 @ B&H)
Power Supply: Corsair RMx 1000 W 80+ Gold Certified Fully Modular ATX Power Supply ($204.99 @ Best Buy)
Case Fan: Corsair LL120RGB LED 43.25 CFM 120 mm Fans 3-Pack ($120.99 @ Best Buy)
Total: $5563.82
Prices include shipping, taxes, and discounts when available
Generated by PCPartPicker 2020-07-15 11:13 EDT-0400
[1] https://pcpartpicker.com/products/memory/#sort=price&U=4&Z=1...
Edit:
Note that they can't be used with multi-gpu builds because they (purposefully) do not have a blower configuration. Unless you can source 2080ti blowers which have the same layout or do a water cooling build it will cause thermal throttling.
Unfortunately the best case for high-mem use-cases is to just rent from GCP.
"NVIDIA TITAN RTX NVLink Bridge
The TITAN RTX NVLink™ bridge connects two TITAN RTX cards together over a 100 GB/s interface. The result is an effective doubling of memory capacity to 48 GB, so that you can train neural networks faster, process even larger datasets, and work with some of the biggest rendering models."
I think only model bigger than gpu mem is where you really wish for nvlink on v100s.
I usually drop a 750 watt 80+ Gold into most of my builds, even though a 500 watt or even a 450 watt would be sufficient with a single GPU, and have no plans for a second GPU.
Efficiency usually starts tapering off below 50% and below 30% it falls off a cliff - however, that just means instead of an ideal 10W power consumption you're actually pulling 30W or something like that, it is usually not a big deal in absolute terms.
(there are also some exceptions, some of the platium/titanium PSUs actually can hold pretty decent efficiencies right down into the basement.)
750W is a good "standard" recommendation, that's enough for any one GPU on the market.
The rule of thumb is really more to guide people not to buy 1600W or 2000W monster PSUs just because "bigger number is better!".
(Although those giant PSUs do have the advantage that they can often run completely passively under load, they won't kick fans on until 50% or 60% load, which for a 1600W PSU means you can comfortably run a high-end GPU and a high-end CPU completely passively.)
No, you're not wrong. Not sure how what I wrote conflicts with that. I don't think it does.
That's what tripped me up, I think. Even without a second GPU, 80+ Gold or better is a good choice.
Your next sentence makes sense though, 1KW PSU's are usually overkill, even if you like to oversize your PSU like I do.
The selection was "gold", specifically. And that's not as bad as it might be, but titanium is better across the board and much better at low load. A titanium supply is more efficient at 20% load than a gold supply at 50%, for instance.
If you're over-sizing your power supply by ~60% (as is the case here) then this is significant.
But, you build enough "rigs", you learn not to skimp on certain components like PSU's, Cases and Motherboards... which is normally where new builders cut corners.
1. https://en.wikipedia.org/wiki/80_Plus#Efficiency_level_certi...
2. https://www.jonnyguru.com/blog/2015/10/25/corsair-rm1000x-10...
tbh this is kind of one of the ideal use-cases for Epyc. And with the way AMD has set up their pricing, it's actually no longer cheaper to use the workstation processors, in some situations it's significantly more expensive, they are really ripping you for the clock speed, and removing a bunch of other features in the process (RDIMM/LRDIMM support, etc). I strongly encourage everyone doing homelab and home ML rigs and similar stuff to really think about whether they want Threadripper, bearing in mind that threadripper is often more expensive than Epyc. It's no longer an obvious choice that server processors are for servers and home users can only afford workstation, it is the other way around.
AMD offers some low-core-count single-socket Epycs that are ideal for "lighting up the platform" tasks like this. Like, 7232P is a $450 processor and the 7402P is $1150. And they don't offer anything like that on Threadripper. They clock slower, sure, but they're not really using the CPU anyway. And that gets you a full 128 PCIe lanes, octochannel memory and RDIMM/LRDIMM support so they can stack in the memory.
If they want to game on it in their spare time then sure, Threadripper is probably the way to go.
I can see the lanes not being an issue for the next gen Titans, that will likely use pcie 4.0, but that is months away.
Asking as someone outside the ML field.
As for RAM, only you can know how big your datasets are. But if you're training models on GPUs the bottleneck is almost certainly going to be GPU RAM, not system RAM.
There is a big delay moving memory from ram to vram to run a task on the gpu, so much so that you'd be better off running the task on the cpu if you can't fit it all in the gpu, or are very clever in how data is buffered, which isn't an option for neural networks. Because of this, the pci-e lane is not saturated except when first sending the data to vram. PCI-E 3.0 x8 runs at 7880MB/s, so if your gpu has 16gb of vram, the difference between x8 and x16 is 1 second, when a task can typically take 8+ hours to complete.
Perhaps for the actual ML part, yes, but a ton of work must be done first to organize and filter the data, which is where all those cores would come in handy.
Even for cheaper builds for non-ML workstations I would still only use Platinum and nothing less. I've been told Titanium is excessive but I mean I leave these things on for a while and power is expensive.
For the DIY enthusiast or the WFH researcher, also the amount of heat involved can be a considerable cost in cooling or utility cost which varies substantially by floor of a building. It's probably not good, but not that bad to aircool this many GPUs as I've done in the past but it definitely means I'm paying a lot for A/C in the summer but almost nothing in the winter.
I think Smerity even said he heated his small bedroom through the San Francisco winter off of one GPU while researching YOLO.
Point: These things get hot. They require a lot of electricity. You should be concerned about a good PSU even for smaller builds. My energy cost for a 6GPU rig ran me about 1/3 of my total rent for a small apartment. That's electricity BEFORE I calculated my A/C bill which was separate and also substantial. My landlord hates me because I initially talked him into including it with my rent.
All in all, it still makes sense to keep investing in local workstations, on-premises builds. No security concerns about a cloud, no futzing around with integrated notebooks, you own it you control it, and the price point up front is extremely attractive compared to base rates for cloud computing even on specialized hardware like a TPU.
The numbers I come up with for batches still have a wide gap of several thousand USD most of the time, and then there's how much time it takes and how likely their service breaks.
So kudos for the person who put in the effort to put this together and share. Any and all efforts towards making ML/DS affordable and DIY rises the tide for all boats.
Question to the audience: Does anyone build GPU rigs like this for cryptocurrency anymore? I was only able to build a workstation once the price for GPU cards crashed.
This guy has a very oversized Gold power supply, the efficiency would be ~92% with gold vs 94% with platinum. Maybe a smaller titanium one would be a better overall choice I guess.
Conversely, this is not a build for overclocking per se. However I think it's a safe assumption we are running at capacity of over 90% for multiple days or weeks even. If batches run into over a month, probably time to get a server rack instead?
It is worthwhile to note you're not going to save any money in efficiency for computation you don't use.
He must not have needed much heat, since a huge GPU would still be 1/6th of a space heater.
For my needs the most intensive thing I do is compile some large programs or encode some large video, which you can get a computer for that for like $800.
I explained the technical details in the sub comment.
I was looking at buying one of these Titan cards a few weeks back but then nvidia announced the next gen processors were coming out so have decided to wait until they refresh the 2 year old titan line instead of paying full prices for an almost out of date card.
When training models for the object detection, the current algo we're using isn't focused on memory efficiency. So the 8gb card we currently use to train models is unable to process images at the correct resolution. We have to down scale about half to get it to fit. With the Titan RTX you get 22gb which is enough.
On another note, the titan cards aren't the same as the normal geforce cards. Nvidia have gone to great lengths to ensure product differentiation so they can charge power users with business budgets more than people sitting at home playing games. One of the good things about the titan cards is they have a dual memory controller so you can write and read at the same time which improves your fill rate.
This can be as simple as identifying when a customer will end their service with a business, as there might be a pattern before previous customers have left, predicting when new customers are going to leave, and giving them a coupon or similar right before they would otherwise leave. This problem is called customer churn.
It can be as complex as identifying when hardware will fail ahead of time, or even bio-ware. For example, I did a project that predicted when people were falling into depression before they could tell they were with a high accuracy rate. I also predicted other future medical issues ahead of time, like the probability an elderly person is going to fall over within the next handful of days.
On the business side there are a lot of use cases for ML, but it falls more into analytics than engineering, as it's about predictive insight.
I really wouldn't build one now unless I had to.
I don't have a ML box yet, let alone a load of servers, but I am contemplating assembling my first rig for ML. If I buy (perhaps secondhand) some GPU, do I risk the thing refusing to work if it incorrectly thinks I'm a server farm?
I have no idea how this could work, or is it just limited to 1 GPU per box? or the proprietary driver phones home? or certain CPU / mobo chipsets are detected and it refuses to run, even if its your only box?
Are you sure? The last time I checked the situation with NVLink memory pooling with 2080ti cards was very unclear.
Looking at something like https://lambdalabs.com/deep-learning/gpu-benchmarks or https://github.com/tensorpack/benchmarks/tree/master/other-w..., multi-gpu scaling on 2080tis seems pretty darn close to linear. Plus, there are benefits to having more than one accelerator handy on a local workstation. For one, it's much easier to have multiple experiments running simultaneously or to run parallel training (e.g. hyperparameter search or RL episodes). Given that only the uber-expensive enterprise cards have proper virtualization/time sharing, trying this workflow on a Titan RTX will most likely be suboptimal unless you always run models that can make use of most of the memory and compute (no RNNs, no Neural ODEs, etc.)
So far the best I have identified is Intel NUC8 + nVidia GPU via Thunderbird. But it is still $1000 at least by the time you have it all together.
NB: I know lots of people will say, just do it with cloud, but I work in a setting where much of my data cannot be put in the cloud, and also where the cost structure of funding well allows for fixed capital expenditure but not variable cloud costs.
By the way, $5k ML workstation is still on the cheaper end of the spectrum. An 8x A100 machine will set you back at least $100k. And even that won't be enough to finetune GPT-3.
GPU via Tunderbolt looks like most expensive way.
Holy moly is that a big heatsink.
But also depends on where in the world you live and what the climate is like.
For the normal desktop cpus, I like the Dark Rock better. Cooling is the same as noctua but was quieter for me.
I should note it's WAY louder than I thought it'd be and it's made getting to my NVME drives a bit difficult.
I didn't do much research beyond "best analog HSF vs watercoolers" and when it showed up I couldn't believe how big it was.
Honestly the best bit though was the piece of mind that I don't have liquid inside my computer anymore. The idea of the pump leaking down onto my GPUs when I'm not around was stressful.
Obviously a homebuilt watercooling rig is going to be way better than AIOs, but it requires extensive maintenance. It's water. Stuff will grow in there. It evaporates. Joints loosen over time and could end up shooting water all over your $6000 workstation.
To me, it's not worth it. Threadripper requires a big heatsink. So be it.
Water cooled PCs do pretty look cool though.
Converting water to steam takes about (from the top of my head) 5x more energy than it takes to increase it by one degree. I imagine the coolant used in these heat sinks is at least this efficient, if not more, and is something that evaporates well below 100°C.
If heat sinks were like using cars and roads to transport people. A water cooled setup would be like expanding the size of the roads to increase the number of cars that can transport people. While those heat pipe coolers are like replacing cars with trains, so you can do a lot more the same amount of space.
I think this greatly underestimates the potential of evaporative cooling. The "specific heat of water" is 1 calorie/gram °C --- that is, one calorie can heat one gram of water by one degree Celsius if no phase change is involved. The "heat of vaporization of water" is more than 500 calorie/gram at 100°C. That is, the energy necessary to convert a given quantity of liquid water from just below boiling to steam is not 5x the energy necessary to raise that amount of water 1°C, but 500x! You are probably remembering that this is 5x the energy necessary to take liquid water all the way from 0°C (almost frozen) to 100°C (almost boiling).
The pump does not need to have air flow to do its job and airflow is what carries noise. The water just has to be able to flow through the loop, pretty much any rate of flow will be pushing enough water to move heat away. If you hold your hand to a water block with the pump turned off, let it warm up and turn it on, it feels like the water is going over your hand. The temperature change is immediate. If you turn off the pump while the CPU is running, the temperature will rise a few degrees each second. Once the pump goes on the temperature drops down immediately.
There is no reason that the liquid should evaporate unless there are leaks and joints don't necessarily loosen over time. Also leaks are unlikely to manifest in something bursting and water spraying, even if you had enough pressure from your pump, which you don't (a pump can probably only pump water four or five feet higher than the pump itself, even if it is more powerful than the loop needs).
I have never had to maintain or fix something after it is together. Use vinyl tubing from the hardware store over the barbs and use small pipe clamps on top of them.
Was fine. I fixed it, refilled it and posted it back! Still running 5 years later.
But I do go a little more paranoid and insist on more expensive silicone tubing(some $$$). Fits nicer, feels nicer, less concern over perishing.
You can have higher tech radiators with better fins along with better fans and airflow modelling etc of course.
But it is always better if there is more of it.
Building an open water loop would be awesome for the GPU cooling. I just can't justify the expense.
Air cooling excels mostly everywhere else.
Neither can cool your system to below ambient room temperature.
Ultimately, you want more surface area to blow air across to cool things. "Water cooling" is really "water transports heat to fins for cooling". Just a bigger version of the heatpipe really.
If the fins are not better suited to airflow cooling then it won't help you.
The AIO radiators don't add that much more fin area than a good air cooler if at all. But you can put them somewhere outside the otherwise warm case if that helps. Maybe if you live some cold and your desk is just right you could dangle it out the window?
Also, for the noise conscious, two(often on the 280 AIOs) or more fans "beat frequency" is a thing. Also why I don't like "push-pull" configs on air coolers.
AWS GPU instances are really expensive.
The most cost effective imo is to build a workstation for development and then deploy to AWS spot if you need a cluster.
If you can't use a workstation for whatever reason, then use the new AWS feature to "stop" spot instances and use the spot instance as your workstation while being conscious of the high hourly cost and shutting it down when you're not working.
I did the maths recently and figured out I could put together a machine with a couple of 2080 Ti’s and have it pay for itself in a couple of months.
I’m very seriously considering doing it, especially as I’m the only data scientist, if I had a team I’d be more in favour of going to the effort of setting up cloud-based training jobs etc
We're waiting for the 3000 series to come out which should be a large performance/dollar improvement over the current gen cards due to the smaller transistor size.
When running cpuburn, i get around 65 tdie temp and with gpuburn, the upper card gets to around 86 and the lower one to 81.
i have right now a water cooler for the cpu, 3 inlet fans (bottom back) and 2 outlet fans through the water cooler on top. i was wondering what temperatures I should aim for and what an optimal fan configuration is. i have a couple of fans laying around.
the case is lian li O11 Air and the mobo is a taichi x399.
anyone any tips?
also I would want to use SLI. but then I would have to remove the fans on the gpu. Do I need then to water cool the gpu or what is the solution?
if anyone of the moddibg pros here can help that would be awesome :-)