> Fully open model: open weights + open data + full training details including all data and training recipes
There are equally open, much more useful models out there: https://artificialanalysis.ai/?models=nvidia-nemotron-3-ultr...
That doesn't mean much to the many people I know of who refuse to use a technology that they see as being unethically created using the work of others without compensating them.
I continue to hope that someone will train a "vegan" model on licensed or out-of-copyright data so those people can experience the benefits of this class of technology.
(I compare them to vegans because, like vegans, I think their ethical position is credible and has merit even though I do not choose the same ethical framework for myself.)
There are models where the weights are released and you can run them locally, but one counter-argument I've seen is that these aren't really the models that people are excited about and making extraordinary claims about.
I haven't seen a model trained on that corpus since late 2024: https://simonwillison.net/2024/Dec/5/pleias-llms/ - I may have missed something though.
Common Corpus is commonly used now in pretraining, including by close labs, but rarely as the only source (which would be the actual ethical commitment).
Like most people training efficient models we move toward synthetic pretraining, but still maintaining our committment for data research and releasability. Lead current project is SYNTH, for now based on Wikipedia but we'll generalize with seeds from Common Corpus.
Nothing below that really seems to be good for anything other than training for specific tasks. I have not been impressed by the earlier Apertus 8B model, which doesn't feel like it really responds to nudges.
I am a strong believer in smaller models, so I might try one of these out of curiosity to see if it might do useful things in limited contexts.