That doesn't mean much to the many people I know of who refuse to use a technology that they see as being unethically created using the work of others without compensating them.
I continue to hope that someone will train a "vegan" model on licensed or out-of-copyright data so those people can experience the benefits of this class of technology.
(I compare them to vegans because, like vegans, I think their ethical position is credible and has merit even though I do not choose the same ethical framework for myself.)
I haven't seen a model trained on that corpus since late 2024: https://simonwillison.net/2024/Dec/5/pleias-llms/ - I may have missed something though.
Common Corpus is commonly used now in pretraining, including by close labs, but rarely as the only source (which would be the actual ethical commitment).
Like most people training efficient models we move toward synthetic pretraining, but still maintaining our committment for data research and releasability. Lead current project is SYNTH, for now based on Wikipedia but we'll generalize with seeds from Common Corpus.
There are models where the weights are released and you can run them locally, but one counter-argument I've seen is that these aren't really the models that people are excited about and making extraordinary claims about.