Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling
arxiv.org
arxiv.org
> We observe that even at the first checkpoint (10B tokens) of WRAP training, the average perplexity of the LLM on the Pile is lower than that achieved by pre-training on C4 for 15 checkpoints. This suggests a 15x pre-training speed-up.