One wonders if they're still doing it.
One wonders if they're still doing it.
Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit!
And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway.
Hell, I know that for one "lab" (kindof AI lab) since 2020 or so has determined wikipedia quality is dropping fast. It was already dropping slowly before that, but now it's getting bad.
Anna's Archive didn't exist in 2019.
It is almost like we need a "for use for training" agreement across the board. This would not fix the current issues (at least without substantial work), but going forward would allow for creators (or publishers/rights holders) such as this to designate a work as crawl-able for AI. A robots.txt just for Claude.
I always wonder why y'all feel the need for these impressive mental gymnastics. You can use the models /and/ think they are trained unethically. Living through the ambiguity without abandoning your ideals completely is a valuable skill these days.
Instead, the ongoing lawsuits focus on the idea that AI training involves making additional copies, for which they would need a copyright license instead of just one legal copy.
Also, as one more example, I find it hard to believe that their models could generate 'Studio Ghibli' style images without training on the movies. There is no licensing deal between them.
I think the real issues here are two-fold:
Firstly, Copyright is very ill equipped to handle these cases. Just because the model is tuned not to output the exact training data does not mean that compressing mostly-copyrighted datasets into a proprietary model is ethical, fair or /should/ be allowed, simply because they might destroy entire livelihoods. If you take those copyrighted works away you are left with, in OpenAIs own words, a cute little experiment.
Secondly, there is absolutely no transparency. Datasets are easily deleted and its impossible to tell what the models have been trained on, especially after fine tuning. Moreover, only the biggest most successfull works would be easily identifiable without the fine tuned model. Once again, sticking it to the little man.