All You Need Is 4x 4090 GPUs to Train Your Own Model
sabareesh.com
sabareesh.com
The best build I have seen so far had 6x4090's. Video: https://www.youtube.com/watch?v=C548PLVwjHA
Specifications
- GPU Accelerator - 6 x 24GB NVIDIA GeForce RTX 4090
- Processor - Intel Xeon W7-3465X, 28C/56T, 2.5GHz - 4.8GHz
- Memory - 256GB (8x32GB) DDR5 ECC 4800MHz
- System Drive - 2TB Samsung 980 PRO NVMe PCIe 4.0 M.2 SSD
- Storage Drive - 4TB Samsung 870 EVO SSD
- Operating System - Ubuntu 20.04
An interesting choice to go with 256GB of DDR5 ECC; if spending so much on the 6x4090's, might as well try to hit 1 TB of RAM as well.The cost of this... not even sure. Astronomical.
Edit: could also just be more-money-than-sense. Never discount stupidity.
The 4090 will likely maintain 50% of its current value due to its memory capacity over the next 12-18 months.
CapEx vs OpEx is a thing even if you are not a business…
The last paragraphs fell totally like AI.
Anyway I'd like a follow up on the curating, cleaning and training part which is far more interesting than how to select hardware which we've been doing for over 25 years.
This jumps right out as written by AI, these features have nothing to do with training LLMs.
Nowadays that's not a lot: a single H100 that you can now rent has 80 GB VRAM, and doesn't have the technical overhead of handling work across GPUs.
It's a rabbit hole I stay away from for pragmatic reasons.
Overall, it's something I've seen very often on social media and less technical articles about LLMs. OpenAI would fall into the "almost" category.
I wasn't nitpicking. It is a HUGE differentiation, and I pointed it out specifically because people pick up on terminology so people who might not know better will go forward and just drop in the more super duper hyperparameter, not realizing that it makes them look like they don't know what they're talking about. As I said in the other post, no one who knows anything uses them interchangeably. It is just completely wrong.
I appreciate being corrected, but you are the one who asked for my opinion based on my extensive time in AI, you can choose to believe it or not.
I see a lot of guides but most focus on getting the toolchain up and running, and not much talk about what kind of dataset do I need to do a good fine tuning.
Any pointers appreciated.
I guess the main question is, do you just prepare training data as if you were training from scratch, or is there some particularities to finetuning that should be considered?
Llama 3.2 Vision is very strictly trained to output a summary at the end which I find difficult to get it stop doing for example.
Another one is that when given a math problem and asked to generate some code that computes the result, most models outputs code fine but insists on doing calculations themselves even if the prompt explicitly say they shouldn't. As expected, sometimes these intermediate calculations are incorrect and hence I don't want the LLM to do that when the produced code would handle it perfectly. If the input prompt contains "four times five" I want the model to generate "4 * 5" rather than "20", consistently.
I've been curious to see if I could tune them to adhere better to the kind of prompts I would be giving.
For LLama 3.2 Vision I've also been curios if I can get it to focus on different details when asked to describe certain images. In many cases it is great but sometimes misses some key aspects.
As for the input training material, that's what I'm trying to figure out what I need. I feel a lot of the guides are like that "how to draw an owl" meme[1], leaving out some crucial aspects of the whole process. Obviously I need input prompts and expected answers, but how many, how much variation on each example, and do I need to include data it was already trained on to avoid overfitting or something like that? None of the guides I've found so far touch on these aspects.
for one, "full" gpu utilization, one or many, remains an open topic in training workflows. spending efforts towards that, while renting from cloud, is a more accessible and fruitful to me than to finetune for marginal improvements.
this course was a nice source of inspiration - https://efficientml.ai/ - and i highly recommend looking into this to see what to do next with whatever hardware you have to work with.
That just doesn't inspire a lot of confidence in those risers, so now I'm contemplating mcio risers.
His power supplies are 2x1500 Watt. That puts it at 3KW max which is more than a 20A circuit can provide (2400W).
The standard outlet is typically rated at 15 amps or 1800W. And the 15A breaker is on one circuit. You can get 20A circuits but they need to be wired for it, and replacing the breaker won't cut it.
Assuming his GPU is ~450W (his number) and power supplies are 80% efficient, well that means he's pulling close to ~2400 watts which is super close to the limit of a 20A circuit.
4 * 450 / 0.80 efficiency = 2250W.
That doesn't include the power consumed by the CPU or mother board or other things on that circuit. But a 170W CPU would easily push this over 2400W provided by a 20A circuit.
It's over current that causes fires.
Commercial setups are not appropriate for typical 15 amp circuit loads.
You can also train your own model even without GPUs. Just depends on parameter size.
HN loves it some Apple
And I say this at the risk of being called pedantic, but a cluster of Mac minis would have zero VRAM.
g Tesla p40 llm reddit
[1]: https://www.pugetsystems.com/labs/articles/llm-inference-con... (8GB model tested but it has same bus width and overall bandwidth as 16GB model)
[2]: https://www.reddit.com/r/LocalLLaMA/comments/1b5uwr4/some_gr...
[3]: https://www.reddit.com/r/LocalLLaMA/comments/178gkr0/perform...
4090s are too small for training and you'll have to write your own suboptimal batching.
Unless you value the learning, it'd be better to rent GPUs in the cloud for training.
This might pull you down a path towards distilling and quantizing models, for instance.
If you want on-prem, wait a few months. The supply of 5000 series (probably announced at CES in a few days) should push more 4000 on the market and, maybe, for a bit, over-supply and push the price down.
Nvidia stopped manufacturing the 4000 a few months ago because they don't have endless factories. Those resources were reallocated to 5000 series and thus pushed the price for the 4000 up to the ridiculous place it is now (about $2,000 on ebay)
I think the current appetite for crypto and ai is big enough to consume all 4000 and 5000 series cards to a point of scarcity (even 3090s are still fetching about $1000) but there should be a window where things aren't crazy expensive coming up.
There's no evidence supply will continually outstrip demand unless something unusual happens.
It's probably somewhere between 12months-never depending on how the market shakes out. Maybe 2 years is a good idea ... really, if power is cheap/free and the machine is on and idle then it's free money - that's the way to look at it.
https://cloud.vast.ai/host/setup
There's a lot of competition in the "airbnb gpu" so if you don't like us, the number is around 12 or so globally. We're probably either #2 or #3. Companies don't really disclose these things so it's hard to know.
Some people probably list on more than one platform. There may be some host management software somewhere that helps with that. I haven't actually checked.
I'd be happy to talk more about these privately. Some are better than others and I've got no interest posting less than charitable things about our competitors publicly, regardless of how accurate I think it is. My email is in my profile.
We aim at $1200/y for 3090, so around a year given descent electricity prices.
Highly recommend setting a lower power limit (usually 250W for 3090).
The industry is full of effectively "imitation companies" right now. For instance, runpod, quickpod, simplepod and clore are the ones cloning us at vast right now.
We see them in our discord, they try to snipe away customers, get in our comment threads on reddit and twitter with self-promotes, clone our features ... this is the ferocious wild west days of this industry. I've even gotten personal emails from a few who I guess scanned their database looking for registration addresses from other companies in the space.
There's even companies like primeintellect which are trying to become the market of markets - but they have their own program - it's clearly a play to snipe other customers by funneling them through some interface where they'll eventually push out the other companies and promote their own instances.
Then there's interesting insider hype players with their own infra like sfcompute who are trying to pretend like they invented interruptible instances and somehow get a bunch of people treating them like they're innovators. The resellable contracts they talk about are a pretty common feature and especially from the host's programmatic command line controller, it's just usually tucked deep in the documentation. They're doing effectively a re-prioritization play.
I guess my angle is "highest integrity possible". It's certainly a gamble - scammy companies sometimes capture a market then become unscammy - I'll hold my tongue but there's plenty of examples.
It's interesting times.
I guess what I’m missing is, what’s scammy about them?
even in the web3 space, AI gpu compute markets are oversaturated
but why is an end user supposed to case about the user acquisition strategy?
if they’re cheaper, more profitable for the gpu owner, or solving a need better, that’s all that matters
You can multi sell a machine, use qemu to lie about the hardware, have hidden fees... there's a bunch of hustle
> AI gpu compute markets are oversaturated
This is not the case. We see a moving average of over 90% utilization of our network. There's a lot of players, but the demand is outstripping supply
> why is an end user supposed to case about the user acquisition strategy?
Well hn is founder/insider talk but for a more direct answer, more legit institutions get higher retention and easier customers.
We're a two sided marketplace so we need to create a platform where people see integrity.
There's also the hypocrisy of complaining about competitors jumping in on "their threads" in a comment on a competitor thread.
Yes, this comment of yours is highly unethical.
But pricing is okay-ish, have a look at Geohot's Tinybox for turnkey solutions.
In Deep Learning it depends on your sharding strategy.
I’m glad to know