* I'm not talking about Roman-era Latin, Latin was the scholarly/international language for Europe for a long long time. AFAIK, most European Latin literature is untranslated...
It's hard to believe that a comparable amount of text had been created, even across all languages, in the entire history of humanity up to the 17th century.
https://huggingface.co/datasets/roneneldan/TinyStories
Also note
https://arxiv.org/abs/2309.05463 (which is larger - obviously size still does contribute to performance)
"RomeGPT" is next on my list of Monad successors and to give you a general idea, we have on the order of tens of millions of words in classical Latin (and biggest source will… Augustine). There was a BERT Latin project that was able to collect roughly 500 million words in all with mostly early modern and modern Latin.
In comparison I'm currently part of a project to pretrain a French model and we need… 140 billion words.
It's like, when you use prompts like : "You're an helpful assistant", is it believing, pretending, or beleiving to be pretending to be an helpful assistant?
It's as funny as disconcerting to see intent and will attributed to probabiltities. Feels sometimes like we're close to making a religion out of this. History of the human race, i guess (https://www.youtube.com/watch?v=xuCn8ux2gbs for the ref).
There are various ways modern knowledge may have "snuck in" through data contamination. If we really want to know what a Shakespearean chatbot would have looked like, we need to cap the training data.
Maybe in the future, using better mechanistic interpretability we can get around this, but not right now.
From the model description
> OpenHermes was trained on 900,000 entries of primarily GPT-4 generated data, from open datasets across the AI landscape.