The problem these AI companies have is they live in a glass house and they can’t throw IP rocks around without breaking their own “your content is our training data” foundation.
They only reason I can think of that Google doesn’t go after OpenAI for scraping YouTube is then they’d put themselves in the same crosshairs, and may set a precedent they’d also be bound by.
Given the model is “on the web” I have the same rights as Mistral to use anything online however I want without regard for IP, right?
Utter absurdity.
Not necessarily. You consented to people reading your code and learning from it when you posted it on Github. Whether or not there's an issue with AI doing the same remains to be settled. It certainly isn't clear cut that separate consent would be required.
They have taken my code and now are dictating how I can use their derived work.
Personally I think these tools are useful, but if the data comes from the commons the model should also belong to the commons. This is just another attempt to gain private benefit from public work.
There are legal issues to be resolved, and there is an explosion of lawsuits already, but the fact pattern is simple and applies to nearly all closed-source AI companies.
[1] https://huggingface.co/replit/replit-code-v1-3b
It'll be a court fight to determine which. Worse, it will be a court fight that plays out in a bunch of different countries and they probably won't all come to the same conclusion. It's unlikely the two licenses have a different effect here though. Either they both forbid it, or neither had the power to forbid it in the first place.
You can profit from GPL / AGPL code but just also make all your source code open source and available for everyone to see.
And if I never posted my code to github, but someone else did? What if someone had posted proprietary code they had no rights to to github at the same time the scraper bots were trawling it? A few years ago some Windows source code was leaked onto Github - did Microsoft consent then?
Has anyone else done this?
I think SV is just dead set on killing the golden goose of open source and the web by extracting as much as possible with no regard for the wasteland left behind.
You are not allowed to reproduce Mistral's works (beyond the usual Fair Use allowances).
Nor is Mistral entitled to reproduce your works (unless you have licensed as such).
If it does, you can sue for copyright infringement.
The point is: nobody knows and the AI companies are getting well ahead of the law.
But what part about what I said do you believe to be undecided?
That a human can learn without violating copyright? That a machine can learn without violating copyright?
Perhaps if all rules were written in stone, and clear of ambiguity, we would not need judges or the legal process. But that’s not how any of this works.
There is some overlap, but if you think you've found some undiscovered loophole in centuries of copyright law, you're mistaken.
It will be the smartphone patent wars all over again with hundreds of lawsuits against big tech and AI companies.
We are already past the 'fair use' excuses at this point especially when OpenAI is slowly striking deals with news companies to train on their content (with their permission) and with intent of commercializing the model.
Which is a great thing.
It's like killing Caesar. As long as we all stab him, everyone is guilty and no one can prosecute us.
While you might call it absurd, I feel like these glass houses are why we've seen so much rapid progress with AI recently.
The library's copyright is intact, as normal, and they can control who uses it and how just like any other software.
The output of AI systems is not copyrightable, but the systems themselves are, and associated EULAs are valid.
Of course, they can revoke your right to use the software, but if it goes to court, that would be interesting case.
I don’t know why there isn’t more discussion on this point and people just assume there’s an underlying copyright basis to the licensing of weights. As far as I know that isn’t settled at all.