MosaicML claims they trained a 7 billion parameter on 1 trillion tokens with a budget of $200k.
https://www.mosaicml.com/blog/mpt-7b
Does training cost scale linearly with model size and token count? If so, that suggests a lower bound of $600k to train the 13 billion params model. (Still roughly the same magnitude)