AI is as good a time as any to re-assess how we feel on intellectual property, but I don't think we can get any more restrictive without suffocating innovation. Plagiarism is how patent trolls shut down otherwise innovative products, and it's the basis for jacking up non-generic drug prices and forcing developers to pay licensing fees just to use an OS. If anything, our current preconception of plagiarism is too unclear and fragile, bound to be destroyed by whatever the hell AI is, if it even matters in the first place.
It's one thing to make something similar, and there are often more ways than one to do that, but how would one know if AI created something unique or just copied some human's copyrighted work and presented it in a response?
TBH, I have not dug into how it works, but I did ask it to show me how to make something with PouchDB.js and it looked like it was pretty much a copy and paste from their website. That is not really an issue, but it did not attribute that code to Pouchdb.com. And, to be fair, I did not ask where it got it. But it seems to me that if they got it off the PouchDB.com web site it should tell us and provide a link to the original source.
You can't. Similarly, when a human writes original code, we can't be certain that they're not repeating stuff they've seen before too. We don't think of things in terms of license, no human remembers the fast inverse square root function or cocktail sort for the license it had. At no point can I be certain that I'm not unconsciously plagarizing proprietary code from a previous job. Sometimes it's essential (I have to plagiarize "set euo pipefail").
There is indeed some nuance to AI creating supposedly-novel works, but I don't think it's as great as people think. AI makes the process of building a derivative work easier, which is probably frustrating as a license-holder. Many forms of derivative works have legal precedent though, like Wine, DXVK, OpenJDK or the Dolphin emulator - it's an open-source implementation of a proprietary API. You can fly pretty close to the sun without getting burned in many of these cases, as long as you don't violate fair use.
Is it even possible to have a novel thought that isn't somehow dependent on an earlier thought of someone else?
No. Who am I plagiarizing when I feel love, hurt, enjoy taste, procreate? So no is the clear answer to that question.
So if you make your model through slow thinking and discovery chatgpt just builds it to a billion words.
The result is the same
A LLM (that hasn't been trained on any data regarding this topic) is unlikely to demonstrate the same bouba/kiki effect seen among humans. https://en.wikipedia.org/wiki/Bouba/kiki_effect
People continue to claim LLMs are just parrots. But then there is never a solid rebuttal to how humans are different.
My point was that it should be.
The training is on images and text, and the output is text. But nothing constrains the model in the middle of those two ends to only model based on language features, as opposed to world features discoverable from the language and images.
To oversimplify a bit, it's a next-word-in-the-sentence prediction engine.
/s
GPTs are similar. They can tell you that the sky is blue and they can tell you why the sky is blue, but they can't translate that to a hypothesis about how the wavelengths of light also influence the properties of lasers. If they've ingested information about these different topics, they can make the connection, but otherwise they can't.
Turns out that when you ask LLMs to predict the next word in the sentence, and then train billions of parameters on billions of sentences, they realize the best way to improve at predicting the next word is to understand why the world works the way it does.