that sucks for using AI right now, but it's only a question of time until you get huge models trained on CC0 data imo
that sucks for using AI right now, but it's only a question of time until you get huge models trained on CC0 data imo
As for your dream of CC0 models, it’s a dream because you’d have to be asleep to believe it. I don’t mean to phrase it so harshly, but there are so many reasons that can’t work. The main one is that there isn’t enough non-copyrighted data to train any competitive model. The competitive models have only recently gotten good enough to just barely be usable, and far more than 90% of their training data was sourced from people they certainly didn’t get usage rights from.
I deleted a paragraph ranting about usage rights. Suffice to say, Stallman’s “Right to Read” becomes more prescient with each passing day.
There's Mitsua Diffusion One [0], which doesn't produce incredible results, but it's a start and they're planning on adding more data, including opt-in work from artists.
PIXART-alpha [1] was trained on only 25 million images, and has excellent and competitive results. This could pair well with Fondant AI's 25 million Creative Commons-only dataset [2] (not all CC0, but a sizeable amount).
I don't think it's as far away as you think it is!
[0]: https://huggingface.co/Mitsua/mitsua-diffusion-one
[1]: https://pixart-alpha.github.io/
[2]: https://huggingface.co/datasets/fondant-ai/fondant-cc-25m