https://github.com/allenai/OLMoE
LLaMA 2: https://arxiv.org/pdf/2307.09288
LLaMA 3: https://arxiv.org/pdf/2407.21783
But the original poster said being glad «they even trained it on open datasets», not glad "they told you what they trained it on".
I do take personal affront when they dump models without specifying their datasets. They're just polluting the information space at that point.
There are two types of "training on copyright material"
1.) Training on material that is copyrighted and behind a paywall, but you circumvent the paywall. This is unambiguously illegal, as the material is pay to view.
2.)Training on material that is copyrighted, but free for anyone to consume. This is ambiguous in legality, but right now seems to be leaning in AI training favor - as long as the models don't verbatim share the material.
There is also another point about copyright material being ad-supported, and obviously the AI doesn't view/care about ads. There is a decent case to be made that this is illegal, but then is ad blocking actually theft?
The point is not about any «AMD's fault». It is about "why would it be great to have LLMs trained on limited (open) data".