This inspires few ideas.
This inspires few ideas.
> A 50 million parameter GPT trained on 5 million games of chess learns to play at ~1300 Elo in one day on 4 RTX 3090 GPUs.
And from the paper: https://arxiv.org/abs/2403.15498
> The 25M parameter model took 72 hours to train on one RTX 3090 GPU. The 50M parameter model took 38 hours to train on four RTX 3090 GPUs.
definitely inspiring :)
A1) Yes it can be done with a 4090. (2a).
A2) 2 days. (4d/4e).
B) Up to you, author did both and settled on 50M which they got to 1300 ELO. (note: they also did subsequent work with the same model to increase perf without further training) (2a).
C) Mu: there's no limit as you can page in/out (3a, 3c), and its trivial to store in both RAM and VRAM. 30B fits in memory with 32 GB of RAM or VRAM. (3e)
Starting from:
1. article: https://adamkarvonen.github.io/machine_learning/2024/03/20/c...
2. second sentence links to post on training: https://adamkarvonen.github.io/machine_learning/2024/01/03/c....
2a. "A 50 million parameter GPT trained on 5 million games of chess learns to play at ~1300 Elo in one day on 4 RTX 3090 GPUs"
2b. "The 50M parameter model played at 1300 Elo with 99.8% of its moves being legal within one day of training"
3. re: most parameters 4090 can do:
3a. understanding is there is no limitation, in that, you don't _need_ to have every parameter in memory at all times, either during training or inference.
3b. google "are amount of parameters in llm limited by vram size"
3c. go to /r/LocalLLaMa link: https://www.reddit.com/r/LocalLLaMA/comments/15j0mvm/what_ar... (why? its a favorite of mine, believe its the closest you get to local training ppl talking in open space that's not discord)
3d. understanding in 3a is correct.
3e. "The model must fit in your RAM or VRAM, but you can split the model between them. With 32gb ram you could fit a 30b model"
4. 3090 vs. 4090 training speed
4a. google "rtx 3090 vs. 4090 llm training perf"
4b. wander through top reddit links. useful info, but a little too technical to share in a way that doesn't require explaining a lot.
4c. down the list: lamda labs link from october 2022: https://lambdalabs.com/blog/nvidia-rtx-4090-vs-rtx-3090-deep...
4d. 4090 is roughly 2x the perf on TransformerXL Large, the best match for transformers based large model, i.e. an LLM
4e. training took 4 3090-sdays for author. So 2 days.
> To sample on Mac, uncomment line 21 in sample.py. To train on Mac, rename train_shakespeare_char_mac.py to train_shakespeare_char.py
The `mac` file changed several things - I decided to try running training with the original config file - changing device to mps / compile to false
iter 100: loss 2.0268, time 815.43ms, mfu 3.24%
iter 200: loss 1.8523, time 818.79ms, mfu 3.24%
iter 300: loss 1.7799, time 823.05ms, mfu 3.23%
iter 400: loss 1.6887, time 819.08ms, mfu 3.23%
Training is ~4x slower than the speed reported on the original multi-GPU run: https://wandb.ai/adam-karvonen/chess-gpt-batch/runs/zt5htyl6...Not bad for an M2 studio which is running lots of other workloads at the same time
It's possible MLX has some additional micro optimizations, but in general most people who have tried it out against hand-written MPS based training implementations haven't found great speed ups yet.