3x4090s 1xTesla A100
Lots of fine tuning, attention visualisation, evaluation of embeddings and different embedding generation methods, not just LLMs though I use them a lot for deep nets of many kinds
Both for my day job (hedge fund) and my hobby project https://atomictessellator.com
It’s summer here in NZ and I have these in servers mounted in a freestanding server rack beside my desk, and it is very hot in here XD
You might like this too - this last weekend I finished writing a topological contour GPU shader for visualizing electron density grids, here is a molecule of buckminsterfullerene - the contour generator is running on the GPU, which is how I can keep the framerate so high.
https://drive.google.com/file/d/1xZkXsWDZtpe0H3BKSkdzvjOA3UE...
Thanks for providing such a wonderful resource for future chemists!
- Apple M2 Max 64GB shared RAM
- Apple Metal (GPU), 8 threads
- 1152 iterations (3 epochs), batch size 6, trained over 3 hours 24 minutes
https://www.reddit.com/r/LocalLLaMA/comments/18ujt0n/using_g...
M2 Ultra Mac Studio vs M3 Max Macbook Pro
https://www.apple.com/shop/buy-mac/mac-studio https://www.apple.com/shop/buy-mac/macbook-pro/16-inch
Does that also apply to stuff like AMD APUs? Honestly asking, I have no clue.
Did the math on how much using runpod per day would be, and bought this setup instead.
Using Fully sharded data parallel and bfloat16, I can train a 7b param model very slowly. But that’s fine for only going 2000 steps!
How is that these days? I water-cooled my system with the intention of having a quiet system sometime around 2008, but it just used liquid to transfer heat to a sink further away, which had a noisy fan.
Thermals are great though. But tbh now what I do is put my machine behind my TV and I have the side open and place the pump outside to stand vertically like old systems that had them outside the case lol. That's because if I'm not playing games on it then there's no difference in just sshing into the system because I live in the terminal anyways.
Though cooling 4090s even with water is hard - you need a good flow rate to dissipate 450W. They'll hit near 70C during training if the fans are kept below audible, which works for me. 55-60C with fans at 50%. This is with 2x420mm (3x 140mm fan) rads and one 280mm rad (2x 140mm fans).
The stock heatsinks actually were actually able to cool them just as well, just they where much louder doing so!
Power is the other tricky bit. The computer is living on the dedicated AC circuit for now... 1600W is a lot. When it was on the single outlet circuit in my house, I could see each iteration as the lights dimmed in sync with the power draw surging.
Even fine tuning Mixtral is 4xH100 for 4 days. Which is a ~$200k server currently.
To fully train, not just fine tune a small model, say Llama 2 7b you need over 128GiB of vram, so still multiple GPU territory, likely A100s or H100s.
This is all dependent upon the settings you use, increase the batch size and you will see even more memory utilization.
I believe a lot of people see these models running locally and assume training is similar, but it isn't.
For pretraining, RAM is the least of the issue. Even for smallest llama 2 they used 12 years of A100 GPU time.