> ~$1 million: Cost to train a 13 billion parameter model on 1.4 trillion tokens
Llama paper mentioned 135,168 A100 hours for training 13 billion model on 1 trillion tokens, which means ~$150k for lambdalabs on demand instance.
Llama paper mentioned 135,168 A100 hours for training 13 billion model on 1 trillion tokens, which means ~$150k for lambdalabs on demand instance.
Plus they don't actually have any actually A100s available at the moment (2022-05-17).
CoreWeave is a nice middle ground. You can at least get the A100 machines into a k8s cluster.