> The source code would include everything needed to train that model and reproduce it.
You know these models are trained on internet scrape which contains copyrighted content, so the dataset can't be open sourced. It's either this or bad models.
You know these models are trained on internet scrape which contains copyrighted content, so the dataset can't be open sourced. It's either this or bad models.
I'm not saying "opening models is bad", it's good. However imo it would be nice to have a semantic way to differentiate between those two
They don't get to claim that it's open source just because it would be too hard to actually open source.