FineWeb: Decanting the web for the finest text data at scale
huggingface.co
huggingface.co
Related: I also suspect that this is one reason we get so little information about the exact data used to train Meta's Llama models ("open weights" vs "open source").
In other words are we just overfitting?
It's important to note that the tests that they use appear to be open source, for example https://huggingface.co/datasets/lighteval/mmlu.
Again, I could be totally ignorant on how these things work. (edited to add key words associated with ChatGPT output in order to increase the quality of my comment :))
Likely some good blog posts will come out and get posted here in the coming days / weeks evaluating and summarizing.