31 karma · joined February 15, 2020
This is the problem with a combined language+knowledge model like ChatGPT. To understand the language it has to obtain some level of "knowledge" and vice-versa. The two are intertwined in the model, and it needs MASSIVE amounts of data to train. Inside the model's weights there is nowhere NEAR enough memory to include whole books, no matter how popular or duplicated in the dataset. Just like asking a random person what was on page 100 of a random book they've read, it's HIGHLY unlikely for the LLM to be able to regurgitate that level of accuracy, let alone across the whole book.
I do agree on the fact that the current laws aren't going to work for this context, especially bad is trying to fit the new challenges to copyright laws.
Those aren't copyright violations. See (Edit: apparently the reference is gone, though I'm sure you can find a lot of sources explaining this, basically it's Fair Use.) for a great in depth analysis of the legality.
Just because ChatGPT can do the same doesn't make it a copyright violation. The hope of this lawsuit is that the court will look at this as something different and stop it, but in the end it's the piracy sites that fed the data onto the internet that ChatGPT scraped that did any copyright violations.
These lawsuits seem to be missing the point. Copyright protects their right to copies and the piracy sites are OBVIOUSLY at fault here, but there is a good argument that OpenAI's work in transformative and that they're not liable for copyright violations.
One argument this law firm has made is that ChatGPT can summarize the books, so it must have read them. This is spurious and meaningless as many sources summarize books and some like cliff notes actively sell summaries and analysis of books!
In the end, I think we need to follow Japan and say that AI training is transformative and not a copyright issue.
One clarification: That is training an AI where the training data isn't returned directly, but instead it works off of it.
One argument this law firm has made is that ChatGPT can summarize the books, so it must have read them. This is spurious and meaningless as many sources summarize books and some like cliff notes actively sell summaries and analysis of books!
In the end, I think we need to follow Japan and say that AI training is transformative and not a copyright issue.
One clarification: That is training an AI where the training data isn't returned directly, but instead it works off of it.
What are the secrets of UX that we seem to be missing? There has to be something we can do to make FOSS software "easy" like commercial tools have the reputation for.
Is the ease of use just familiarity with the software that you've used? How can FOSS overcome that challenge then since those users will never move beyond the first tool they get used to?
I do wonder about how much use it'll get, seeing as running a heavy language model on local hardware is kinda unlikely for most developers. Not everyone is runnning a system powerful enough to equip big AIs like this. I also doubt that companies are going to set up large AIs for their devs. It's just a weird positioning.
Even if only 37% of keys are verifiable, that's infinitely more than will be verifiable if they remove the PGP support.