Building a deep learning rig
samsja.github.io
samsja.github.io
For those talking about breakeven points and cheap cloud compute, you need to factor in the mental difference it makes running a test locally (which feels free) vs setting up a server and knowing you're paying per hour it's running. Even if the cost is low, I do different kinds of experiments knowing I'm not 'wasting money' every minute the GPU sits idle. Once something is working, then sure scaling up on cheap cloud compute makes sense. But it's really, really nice having local compute to get to that state.
With a local setup, I often think, "Might as well run that weird xyz experiment over night" (instead of idling) On a cloud setup, the opposite is often the case: "Do I really need that experiment or can I shut down the sever to save money?". Makes a huge difference over longer periods.
For companies or if you just want to try a bit, then the cloud is a good option, but for (Ph.D.) researchers, etc., the frictionless local system is quite powerful.
but if you're just goofing around and not planning to create anything production worthy, it's a great deal.
vast.ai is basically a clearinghouse. they are not doing some VC subsidy thing
in general, community clouds are not suitable for commercial use.
Are you factoring in the varying power usage in that electricity price?
The electricity cost of operating locally will vary depending on the actual system usage. When idle, it should be much cheaper. Whereas in cloud hosts you pay the same price whether the system is in use or not.
Plus with cloud hosts reliability is not guaranteed. Especially with vast.ai, where you're renting other people's home infrastructure. You might get good bandwidth and availability on one host, but when that host disappears, you should hope that you did a backup, which vast.ai charges for separately, and if so, you need to spend time restoring the backup to another, hopefully equally reliable host, which can take hours depending on the amount of data and bandwidth.
I recently built an AI rig and went with 2x3090s, and am very happy with the setup. I evaluated vast.ai beforehand, and my local experience is much better, while my electricity bill is not much higher (also in EU).
Agreed on reliability and data transfer, that's a good point.
Out of curiosity, what do you use a 2x3090 rig for? Bulk not time-sensitive inference on down quanted models?
If you're using them for inference, your usage pattern is unpredictable. I could spend hours between having to use it, or minutes. If you shut it down and release it, the host might be gone the next time you want to use it.
> what do you use a 2x3090 rig for? Bulk not time-sensitive inference on down quanted models?
Yeah. I can run 7B models unquantized, ~13-33B at q8, and ~70B at q4, at fairly acceptable speeds (>10tk/s).
i do think that cloud GPUs can cover most of this experimentation/learning need.
This is obviously because their are forced to use high memory cards.
Are there ideal cards for low memory (1-2BN) models? So higher flops/$ on crippled memory
fwiw I find runpod's vast clone significantly better than vast and there isn't really a price premium.
Is there a goto card for low memory (1-2BN) models?
Something with much better flops/$ but purposely crippled with low memory.
- if I have it locally, I'll play with it
- if not, I won't (especially with my data)
- if I have something ready for a long run I may or may not want to send it somewhere (it's not going to be on 3090s for sure if I send it)
- if I have requirement to have something public I'd probably go for per usage with ie [0].
I've had VERY hit-miss results with Vast.ai and I'm convinced people are cheating their evaluation stuff because when the rubber meets the road it's very clear performance isn't what it's claimed to be. Then you still need to be able to actually get them...
Unfortunately my CFO (a.k.a Wife) does not share the same understanding.
(not really, but it is a joke I read someplace and I think it applies to a lot of couples).
Device 0 [NVIDIA GeForce RTX 3060] PCIe GEN 3@16x RX: 0.000 KiB/s TX: 55.66 MiB/s GPU 1837MHz MEM 7300MHz TEMP 43°C FAN 0% POW 43 / 170 W GPU[|| 5%] MEM[|||||||||||||||||||9.769Gi/12.000Gi]
Device 1 [Tesla P40] PCIe GEN 3@16x RX: 977.5 MiB/s TX: 52.73 MiB/s GPU 1303MHz MEM 3615MHz TEMP 22°C FAN N/A% POW 50 / 250 W GPU[||| 9%] MEM[||||||||||||||||||18.888Gi/24.000Gi]
Device 2 [Tesla P40] PCIe GEN 3@16x RX: 164.1 MiB/s TX: 310.5 MiB/s GPU 1303MHz MEM 3615MHz TEMP 32°C FAN N/A% POW 48 / 250 W GPU[|||| 11%] MEM[||||||||||||||||||18.966Gi/24.000Gi]
You can expect a GPU to last 5 years. So for 128 days break even you are only looking at 6.67% utilization. If you are doing training runs, I think you are going to beat it easily.
P.S. coincidentally or not, but shortly after it got mentioned on Hacker News, Best Buy run out of both RTX 4090s and RTX 4080s. They used to top the chart. Turns out at descent utilization they win due to the electricity costs.
[0] https://www.royalgazette.com/general/business/article/202307...
You have to ask yourself if you want to drop that kind of money on consumer GPUs, which launched late 2022. But then again, with that kind of money you are stuck with consumer GPUs either way, unless you want to buy Ada workstation cards for 6k each and those are just 4090s with p2p memory enabled. Hardly worth the premium, if you don't absolutely have to have that.
which means you could build a 4gpu server from normal cases.
most of the 4090 cards are 2-3 slot cards
so you are getting a better practical card for the price
if you are making a mining type rig, then yeah, the extra price is wasting money.
but if you wanted to build a normal machine, the workstation cards are the most reasonable choice for anything more than 2 gpus
Only if you are already live near an airport and you are accustomed to the sounds of the lift off and flying away.
People complain about the "Nvidia tax". I don't like monopolies and I fully support the efforts of AMD, Intel, Apple, anyone to chip away at this.
That said as-is with ROCm you will:
- Absolutely burn hours/days/weeks getting many (most?) things to work at all. If you get it working you need to essentially "freeze" the configuration because an upgrade means do it all over again.
- In the event you get it to work at all you'll realize performance is nowhere near the hardware specs.
- Throw up your hands and go back to CUDA.
Between what it takes to get ROCm to work and the performance issues the Nvidia tax becomes a dividend nearly instantly once you factor in human time, less-than-optimal performance, and opportunity cost.
Nvidia says roughly 30% of their costs are on software. That's what you need to do to deliver something that's actually usable in the real world. With the "Nvidia tax" they're also reaping the benefit of the ~15 years they've been sinking resources into CUDA.
I wonder if it has anything to do with the strategy they used in 3D graphics, where game developers ended up writing for NV drivers in order to maximise performance, and Nvidia abused their market position and used every trick in the book to make AMD cards run poorly. People complained about AMD driver quality, but the actual problem was that they were not NV drivers and AMD couldn't defeat their software moat.
So here we are again, this time with AI. You'd think we'd learnt our lesson but instead people are fooled yet again and instead of understanding the importance of diversity and competition in the lifeblood of their art, myopia and amnesia is the order of the day.
Tinygrad are doing god's work and I won't be giving Nvidia a single fucking cent of my money until the software is hardware neutral and there is real competition.
I find this vague reflexive anti-corpo leftism that seems to have become extremely popular post-2020 really tiresome.
Ah ideology, such a great alternative to actual thinking. Don't investigate or reason, just just blame it on the 'lefties'. Tiresome indeed.
Not sure how me simply stating the obvious makes me a 'lefty'. If you think monopiles, regardless of how they come about, are a good idea, that companies should be allowed to lock up an important market for any reason, then that makes you a corporatist fascist, right? Wow, this mindless name calling is so much fun! I feel like a total genius.
The simple fact is that it is the nature of software, its complexity and dependence on a multitude of fairly arbitrary technical choices makes it a very effective as a moat, even if its not intentional. CUDA, etc is 100% a software compatibility issue, and that's it. There's more than one way to skin a cat but we're stuck with this one. Nvidia isn't interested in interoperability, even though it's critical for the industry in the longer term. I'd wouldn't be either if it was money in my pocket.
The point that is entirely missed here is that we, as a community, are screwing up by not steering the field toward better hardware compatibility, as in anyone being able to produce new hardware. In the rush to improve or or try out the latest model or software we have lost sight of this, and it will be to our great detriment. With the concentration of money in one company we will have a lot less innovation overall. Prices will be higher and resources will be misallocated. Everyone suffers.
It's very possible that AI withers on the vine due to lagging hardware. It's going to need a lot of compute, and maybe a different kind of compute to boot. We may need a million or a billion times what we have to even get close to AGI. But if one company locks that up, and uses that position to squeeze out every dollar from its customers (really, have a look at the almost comical 'upgrades' Nvidia offers in their GPUs other than at the very high end) then it's going to take much longer to progress, and maybe we never get there because some small group of talented maverick researchers were never able to get their hands on the hardware they needed and never produce some critical breakthrough.
No, there is a limit to the mount of handwringing, begging, and crying the community can do that would have forced AMD to take GPGPU computing seriously. CUDA didn't spring out of no where, it's 16 years old and in that time many people have begged AMD to properly support OpenCL or RocM. It's not the communities fault that AMD didn't take this field seriously until it was too late. Seriously, the consumer GPUs don't even get official RocM support, but somehow it's nvidia's fault that AMD didn't care to support RocM.
I'm sure AMD will wake up now that CUDA is a trillion dollar market, but it's unfair to blame users supporting CUDA. nvidia invested in open source for more than a decade now, and there were people who foresaw the current situation and tried to develop more open backends for frameworks like torch. Unfortunately developers don't work for free and nvidia spent the money and AMD did not. It's not users fault that they didn't work, for free, to get tensorflow working on AMD.
Geohot[1] nearly gave up on AMD entirely when their own drivers don't work. This isn't new, AMD is culpable for the current situation, the community didn't end up here due to indifference.
Can we drop the "Nvidia is the only self-interested evil company in existence" schtick?
I'm not being "fooled" by anyone. I've been trying to use ROCm since the initial release six years ago (on Vega at the time). I've spent thousands of dollars on AMD hardware over the years hoping to see progress for myself. I've burned untold amounts of time fighting with ROCm, hoping it's even remotely a viable competitor to CUDA/Nvidia.
Here we are in 2024 and they're still doing braindead stuff like dropping a new ROCm release to support their flagship $1000 consumer card a full year after release...
ROCm 6 looks good? Check the docker containers[0]. Their initial release for ROCm only supported Python 3.9 for some strange reason even though the previous ROCm 5.7 containers were based on Python 3.10. Python 3.10 is more-or-less the minimum for nearly anything out there.
It took them 1.5 months to address this... This is merely one example, spend some time actually working with this and you will find dozens of similar "WTF?!?" bombs all over the place.
I suggest you put your money and time where your mouth is (as I have) to actually try to work with ROCm. You will find that it is nowhere near the point of actually being a viable competitor to CUDA/Nvidia for anyone who's trying to get work done.
> Tinygrad are doing god's work
Tinygrad is packaging hardware with off the shelf components plus substantial markup. There is nothing special about this hardware and they aren't doing anything you couldn't have done in the past year. They have been vocal on calling out AMD but show me their commits to ROCm and I'll agree they are "doing god's work".
We'll save the work being done on their framework for another thread.
What they are doing is all of the hardware engineering work that it takes to build something like this. You're dismissing the amount of time they spent on figuring stuff like this out:
"Beating back all the PCI-E AER errors was hard, as anyone knows who has tried to build a system like this."
Define "hard".
The crypto mining community has had this working for at least half a decade with AMD cards. With Nvidia it's a non-issue. I'd be very, very curious to get more technical details on what new work they did here.
That said, if you think any of this is easy, you're the one who should define that word.
With that said.
Easy: Assembling off the shelf PC components to provide what is fundamentally no different than what gamers/miners build every day. Six cards in a machine and two power supplies is low-end mining. Also see the x8 GPU machines with multiple power supplies that have been around forever. I'm not quite sure why you're arguing this so hard, you're more than familiar with these things.
Hard: Show me something with a BOM. Some manufacturing? PCB? Fab? Anything.
FWIW for someone that is frequently promoting their startup here you come across as pretty antagonistic. I'm not attacking you, just saying that for someone like myself that has been intrigued by what you're working on it gives me pause in terms of what I'd charitably refer to as potential personality/relationship issues.
Everyone has those days, just thought it was worth mentioning.
For mining, the focus is mainly on GPUs, and the specifics like bus speed or other components aren't as critical. You could get by with a basic setup - a $35 CPU, 4GB of RAM, 100meg network, and PXE booting without any local storage. Even older GPUs like the RX470s did the job perfectly until the very end.
But what George is working on is something else entirely. It's not just about the number of GPUs; it's about creating a cohesive system where every component plays its part and is configured correctly. This complexity of tying everything together, is what makes it challenging. George is incredibly talented, and the fact that he's been dedicating himself to the tinybox project for a year now really speaks volumes about the intricacies involved.
Please don't think I'm trying to be confrontational - that's not my intention in the slightest. I appreciate your perspective, but I'm just trying to offer a different angle based on my own experience in this field.
While it might seem that this hardware isn't groundbreaking or that the developments could have been achieved earlier, it's important to recognize the innovation and hard work behind it. This isn't just about putting together existing pieces; it's about creating something that works better as a whole than the sum of its parts.
I'm confident that if we were to talk in person, we'd get along just fine.
It's not that bad. You just copy and paste ~10 bash commands from their the official guide. 7900XTX is now officially supported by AMD. Andrew Ng says it's much better than 1 year ago and isn't as bad as people say.
I have a 3U supermicro server chassis that I put an AM4 motherboard into, but I'm looking at upgrading the Mobo so that I can run ~6 3090s in it. I don't have enough physical PCIE slots/brackets in the chassis (7 expansion slots), so I either need to try to do some complicated liquid cooling setup to make the cards single slot (I don't want to do this), or I need to get a bunch of riser cables and mount the GPU above the chassis. Is there like a JBOD equivalent enclosure for PCIE cards? I don't really think I can run the risers out the back of the case, so I'll likely need to take off/modify the top panel somehow. What I'm picturing in my head is basically a 3U to 6U case conversion, but I'm trying to minimize cost (let's say $200 for the chassis/mount component) as well as not have to cut metal.
They have single-slot GPU waterblocks but would want something like $400 or more each for them individually.
For the chassis, you could try a 4U rosewill like this: https://www.youtube.com/watch?v=ypn0jRHTsrQ, not sure if 6 3090s would fit though. You're probably better off getting a mining chassis, it's easier to setup and cool down, also cheaper, unless you plan on putting them in a server rack.
Am also inspired by embedded developers for the same reason
p.s. the mobo (B450 Steel Legend) already has 2 pcie x16 slots, so the recommendation does not make sense to me.
Though you're right of course that pcie will totally suffice for this case.
I would prefer a tutorial on how to do this.
My box has a Gigabyte B450M, Ryzen 2700X, 32GB RAM, Radeon 6700XT (for gaming/streaming to steam link on Linux), and an "old" Geforce GTX 1650 with a paltry 6GB of RAM for running models on. Currently it works nicely with smaller models on ollama :) and it's been fun to get it set up. Obviously, now that the software is running I could easily swap in a more modern NVidia card with little hassle!
I've also been eyeing the b450 steel legend as a more capable board for expansion than the Gigabyte board, this article gives me some confidence that it is a solid board.
ROMED8U-2T
Are you on the latest bios? I have had good results flashing the "E" bios in the TNT folder based on the FTP site mentioned here
The main benefit is you can shut off nodes entirely when not using them, and then when you turn them back on they just rejoin the cluster.
It also helps managing different types of devices and workloads (tpu vs gpu vs cpu)
I'm not sure what I'd use it for.
2 x RTX4090 workstation guide
You can put two aircooled 4090 in the same ATX case if you do enough research.
https://github.com/eul94458/Memo/blob/main/dual_rtx4090works...
those rigs need pcie riser slots that are also limited.
looks like the primary value is the rig and the cards. they'll need another 1-2k for a thread ripper and then the riser slots.
They also have some ML inference stuff on chip themselves.
inb4 there are no cloud 3090s: yes there are, just not in formal datacenters
Llama definite a bit of a different story though.