So how do I use this? As someone new to the domain.
Is 40 a100s enough though? I am interested in what this would cost.
> When training a 65B-parameter model, our code processes around 380 tokens/sec/GPU on 2048 A100 GPU with 80GB of RAM.[1]
Note that you probably need to budget for double to triple that because things go wrong and it usually takes multiple starts to get a good training run.
Smaller models are cheaper though.