A1) Yes it can be done with a 4090. (2a).
A2) 2 days. (4d/4e).
B) Up to you, author did both and settled on 50M which they got to 1300 ELO. (note: they also did subsequent work with the same model to increase perf without further training) (2a).
C) Mu: there's no limit as you can page in/out (3a, 3c), and its trivial to store in both RAM and VRAM. 30B fits in memory with 32 GB of RAM or VRAM. (3e)
Starting from:
1. article: https://adamkarvonen.github.io/machine_learning/2024/03/20/c...
2. second sentence links to post on training: https://adamkarvonen.github.io/machine_learning/2024/01/03/c....
2a. "A 50 million parameter GPT trained on 5 million games of chess learns to play at ~1300 Elo in one day on 4 RTX 3090 GPUs"
2b. "The 50M parameter model played at 1300 Elo with 99.8% of its moves being legal within one day of training"
3. re: most parameters 4090 can do:
3a. understanding is there is no limitation, in that, you don't _need_ to have every parameter in memory at all times, either during training or inference.
3b. google "are amount of parameters in llm limited by vram size"
3c. go to /r/LocalLLaMa link: https://www.reddit.com/r/LocalLLaMA/comments/15j0mvm/what_ar... (why? its a favorite of mine, believe its the closest you get to local training ppl talking in open space that's not discord)
3d. understanding in 3a is correct.
3e. "The model must fit in your RAM or VRAM, but you can split the model between them. With 32gb ram you could fit a 30b model"
4. 3090 vs. 4090 training speed
4a. google "rtx 3090 vs. 4090 llm training perf"
4b. wander through top reddit links. useful info, but a little too technical to share in a way that doesn't require explaining a lot.
4c. down the list: lamda labs link from october 2022: https://lambdalabs.com/blog/nvidia-rtx-4090-vs-rtx-3090-deep...
4d. 4090 is roughly 2x the perf on TransformerXL Large, the best match for transformers based large model, i.e. an LLM
4e. training took 4 3090-sdays for author. So 2 days.