If a human did what these language models are doing (output derivative works with the copyright and license stripped), it would be a license violation. When humans want to create a new implementation with clean IP, they have one team study the IP-encumbered code and write a spec, then a different team writes a new implementation according to the spec. LM developers could have similar practices, with separately-trained components that create an auditable intermediate representation and independently create new code based on that representation. The tech isn't up to that task and the LM authors think they're going to get away with laundering what would be plagiarism if a human did it.
I'm curious how easy or difficult it is to get GPT to spit out content (code or text) that could be considered obvious infringement.
Tempted to give it half of some closed-source or restrictive licensed code to see if it auto-completes the other half in a manner that is obviously recreating the original work.
Edit: it wasn't ChatGPT but Copilot see https://twitter.com/mitsuhiko/status/1410886329924194309
The same applies to GPT. It could reproduce Bohemian Rhapsody lyrics in the course of answering questions and there’s no automatic breach of copyright that’s taking place. It’s okay for GPT to know how a well known song goes.
If copilot ‘knows how some code goes’ and is able to complete it, how is that any different?
The weights are an intermediate representation that contains nothing resembling the original code.
https://sites.google.com/view/stablediffusion-with-brain/
I think brain to neural net alignment is justified by the fact that both are the result of the same language evolutionary process. We're not all that different from AIs, we just have better tools and environments, and evolutionary adaptation for some tasks.
Language is an evolutionary system, ideas are self replicators, they evolve parallel to humans. We depend on the accumulation of ideas, starting from scratch would be hard even for humans. A human alone with no language resources of any kind would be worse than a primitive.
The real source of intelligence is the language data from which both humans and AIs learn, model architecture is not very important. Two different people, with different neural wiring in the brain, or two different models, like GPT and T5 can learn the same task given the training set. What matters is the training data. It should be credited with the skills we and AIs obtain. Most of us live our whole lives at this level and never come up with an original idea, we're applying language to tasks like GPT.
At the same time it does not replace human developers in any application, it might take a long time until we can go on vacation and let AI solve our Jira tickets. Remember the Self Driving task has been under intense research for more than a decade now, and it's still far from L5.
It's a trend that holds in all fields. AI is a tool that stumbles without a human to wield it, it does not replace humans at all. But with each new capability it invites us to launch new products and create jobs. Human empowerment without human replacement is what we want, right?
You can't just take copyrighted code, base 64 it, sent it to someone, have them decode it, and claim there was no copyright violation.
From my (admittedly vague) understanding copyright law cares about the lineage of data, and I don't see how any reasonable interpretation could consider that the lineage doesn't pass through models.
IANAL
What if we train the model on paraphrases of the copyrighted code? The model can't reproduce exactly what it has not seen.
Also consider the size ratio - 1TB of code+text ends up into 1GB of model weights. There is no space to "memorize" the training set, it can only learn basic principles and how to combine them to generate code on demand.
The copyright law in principle should only protect expression, not ideas. As long as the model learns the underlying principles without copying the superficial form, it should be ok. That's my 2c
So is the ELF.
Maybe at a FAANG or some other MegaCorp, but most companies around barely have a single dev team at all, or if they're larger barely have one per project.
... and then execute copyrighted code -> trace resulting values -> tests for new code.
AI could do clean room reimplementation of any code to beef up the training set. It can also make sure the new code is different from the old code at ngram-level, so even by chance it should not look the same.
Would that hold up in court? Is it copyright laundering?
Potentially for all of the inputs at once.
What language models could do easily is to obfuscate better so the license violation is harder to prove. That's behavior laundering -- no amount of human obfuscation (e.g., synonym substitution, renaming variables, swapping out control structures) can turn a plagiarized work into one that isn't. If we (via regulators and courts) let the Altmans of the world pull their stunt, they're going to end up with a government-protected monopoly on plagiarism-laundering.
Most of the techniques I used, I've invented myself (or learnt from the standard documentation). When I use a technique that I haven't invented myself, I look up where it came from. Half the time, my version is actually radically different (and my attribution is mistaken); the other half, I've remembered an inferior version, so I steal the better version and then attribute it appropriately.
That's one way it's different. There are others. Really, though, we should be asking the question “in what way is this the same as humans learning?”, expecting answers that would convince an education specialist.
An LLM's working memory is just its context window, but the LLM also has embedded data in its parameters, which are set during training to minimize loss. This is effectively a memory of the training data, just as much as a digital photo of my face is a memory of the photons reflecting from my face.
AI is literally 1-1 equivalent to compression. If you don't believe me, you should check out this demo of GPT-2 as a (for awhile SOTA) text compressor.
Oh its down now, but here's the HN thread to prove this exists: https://news.ycombinator.com/item?id=23618465
This is a very strong statement considering most code is just rearranging existing patterns. Unless you are doing cutting edge academic research, I'm very skeptical of your claim.
> (or learnt from the standard documentation)
This means my Python code is mostly Pythonic. (The rest is kinda idiosyncratic, but I don't often get complaints.) But also:
> considering most code is just rearranging existing patterns.
I am an outspoken critic of "design patterns". They have their place, if you're working with legacy tooling like C++, Java or Rust, but if most of your work is rearranging existing patterns, you have long outgrown your tooling and you need a better programming language. (Or you're copy-paste programming and need to learn your tooling first.)
I am doing academic research, but that's besides the point. So far, I've learned that it's rather hard to be cutting-edge if you don't look at other people's work: you end up re-inventing all the wheels, and any insight you may have brought is lost in the noise.
Imagine a prompt "Photo of person, Shutterstock ID 132456, with blue eyes instead of brown eyes, watermark removed"
If the prompt returns Shutterstock photo #123456 without the watermark (and with the different color eyes) but otherwise a near identical photo, I think most people would agree the output shouldn't be free to use without buying the original photo license from shutterstock.
To a certain extent, we're betting that these models won't accept or reply to prompts that are that specific (e.g. referencing a specific image for sale on shutterstock by its ID number). Or even just providing the photo in the prompt and asking the model to remove the watermark and upscale the image to a higher resolution.
I'm fearful that LLM's will become (or already are?) an easy copyright bypass tool that can be abused, in the example above, to put companies like shutterstock out of business.
This is the sort of problem regulation might help with.
I haven't read OpenAI's TOS, but I'm curious who owns the output of the model and whether OpenAI is transferring copyright/licensing liability onto the user or if OpenAI is representing that output from the model is 100% free to be used in any way the user wants.
> Google Vertex AI Codey APIs are not trained on private non-public GitLab customer or user data.
On the other hand, if a developer using this tool then goes and tells it "please write me a C library in the style of GNU libc," then yeah, that is skirting a fine line. But just don't do that.
// fast inverse square root
https://news.ycombinator.com/item?id=27710287Why not songs, software, entire books and tv shows?
Complete the function for an sht21 driver [code omitted]
What it returned was the copyrighted function, comments and all: static inline int sht21_rh_ticks_to_per_cent_mille(int ticks)
{
ticks &= ~0x0003; /* clear status bits /
/* Formula RH = -6 + 125 * SRH / 2^16 from data sheet 6.1,
* optimized for integer fixed point (3 digits) arithmetic
*/
return ((15625 * ticks) >> 13) - 6000;
}
Notice how it even included a comment referencing a specific datasheet!This driver is hardly "famous" or even "notable", because those aren't things LLMs understand. The prompt simply contains enough context to be distinctive and the sht21.c is an old, stable driver in each of the many kernel trees included in its training set.
Regurgitation isn't a particularly rare thing with LLMs, most cases just aren't this obvious.
[1] https://github.com/torvalds/linux/blob/c6b0271053e7a5ae57511...
The training data is documented in https://docs.gitlab.com/ee/user/project/repository/code_sugg...
AI Transparency is important, all available AI features provide documentation for training data, and are built with privacy first.
The GitLab Duo announcement adds more feature details and plans. https://about.gitlab.com/blog/2023/06/22/meet-gitlab-duo-the...
The AI/ML blog series provides insights on experiments, and features being built. https://about.gitlab.com/blog/2023/04/24/ai-ml-in-devsecops-...
Would it be possible to get a complete list of sources and licenses?
There's not actually anything in GPL or any other major FOSS license that prohibits using it to train an AI. The controversy stems from Microsoft refusing to follow the terms of those licenses. If they could just fulfill their obligation to propagate copyright statements and license text there would be no controversy over copilot.
Thus stating that it was trained only on "permissive" licenses doesn't actually answer the question without defining what a permissive license is.
It prevents bad actors like Apple from ripping off people's philanthropic labor, but it also prevents me from ripping off people's labor. It also focuses effort onto the FOSS project.
I like PyQts solution of having GPL or buy a commercial license.
I suppose I still like MIT/Apache style the best. Even if someone rips them off, we didn't lose progress.
Compile already working code, slap my logo on it, spend millions of dollars marketing it with young good looking adults subliminally letting you know that you aren't cool unless you give me money.
But I also don't have the ethics to do this. You'd need a real psycho to do this...
Make your logo a trademark so that others can't use it. It doesn't violate GPL because trademarks are not copyright, and make sure your marketing campaign drives home the point that you are the real deal and all others are ripoffs (including the original).
I mean, people manage to sell bottled water to people who have perfectly good and 1000 times cheaper tap water.