Llama 1.3B Trained on 200B Tokens for Commercial Use
huggingface.co
huggingface.co
> The model was pre-trained for 200B tokens (batch size 2200, sequence length 2048). It was trained on the following data mix:
67% RedPajama Common Crawl
15% C4
4.5% RedPajama GitHub
4.5% RedPajama Wikipedia
4.5% RedPajama Books
2.5% RedPajama Arxiv
2% RedPajama StackExchange
Any benchmarks on model performance?Here is the goal of Chinchilla scaling from https://arxiv.org/pdf/2203.15556.pdf : "In this work, we revisit the question: Given a fixed FLOPs budget, how should one trade-off model size and the number of training tokens? To answer this question, we model the final pre-training loss L(N, D) as a function of the number of model parameters N, and the number of training tokens, D. Since the computational budget C is a deterministic function FLOPs(N, D) of the number of seen training tokens and model parameters, we are interested in minimizing L under the constraint FLOPs(N, D) = C."
But this isn't the only interesting optimization. It can also be useful to have small models that are "overtrained" relative to chinchilla optimality. There are practical reasons to prefer smaller models over big models even at some cost of training expense or pre-training loss, but the chinchilla optimization does not account for this at all. For one thing, the smaller models can fit in the memory of a wider class of devices. For another thing, smaller models incur less computational expense at inference time. Yet another reason is that they will probably have less latency.