Ownership of AI-Generated Code Hotly Disputed
spectrum.ieee.org
spectrum.ieee.org
Is it really feasible? What does “traces attribution” even mean here? It’s not emitting “code”, it’s emitting individual tokens that each were found throughout the input corpus. The “code” is the arrangement of those tokens, but that is determined by the weighting of the whole network, so what can be traced?
Can someone who understands generative ML better than me weigh in on this?
If such a giant list of acknowledgements were to be included, would it even help anyone, in any way? It sounds like pure inefficiency.
This theory of operation is not borne out in reality. It's been clearly displayed that these tools are emitting verbatim copies of existing code (and its comments) in their input.
It's even being seen in image generation, where NatGeo cover images are reproduced in their entirety or where stock photo watermarks are emitted on finished images.
And so, what can be traced back to individual sources? Quite a bit it would seem.
And if you can trace how data moves through the model, you can identify "that output came from these inputs" and add the associated metadata.
Can you cite sources? I’ve heard this claim repeatedly but have yet to see a good example.
https://news.ycombinator.com/item?id=33061707 <- Watermarks (there's been a few linked here on HN)
It doesn't mean the rest of the output existed in the training set. It's perfectly capable of generating a completely novel picture, then slapping a watermark on it.
Stable diffusion is more like a paintshop artist who grabs bits and pieces of other art and melds them together, and less like a painter who creates from their imagination.
I respectfully disagree with this claim. It’s not as-is. It’s remarkably similar. That’s a big difference.
> Stable diffusion is more like a paintshop artist who grabs bits and pieces of other art and melds them together, and less like a painter who creates from their imagination.
I also disagree with this. The uncomfortable truth, imho, is that what Stable Diffusion does is FAR closer to what human artists do than we’d like to admit.
Human artists are perfectly capable of reproducing copyright infringing images. They just generally choose not to for legal/moral reasons.
This is only possible for elements that are repeated hundreds or thousands of times. It's generally considered a failure of training data curation, if it's anything more than watermarks. The AI can learn how to reproduce the watermark, while being literally unable to reproduce the pictures it was on.
I think it’s pretty obvious that if you train a model to reproduce existing work in its entirety, it fails the “sufficiently transformative” test and thus loses legal protection.
But there’s nothing stopping you from re-coding existing implementations of GPL’ed code. I used to do it. And your new code has your own chosen license, even if your ideas came from someone else. Are you sure the same logic shouldn’t apply to models?
(def no (x)
(id x nil))
(def atom (x)
(no (id (type x) ‘pair)))
(def some (x f)
(if (no x) nil
(f (car x)) x
(some (cdr x))))
(def all (x)
(if (no x) t
(f (car x)) (all (cdr x))
nil)))
I don’t even have to pull up bel.bel to know that those are almost perfect replicas. I typed it on my iPad.EDIT: as far as I can tell, the only diff is that all comes before some. https://sep.yimg.com/ty/cdn/paulgraham/bel.bel?t=1595850613&
I could have kept going through most of the implementation.
Perhaps my example isn’t as impressive as the model’s capability, but operationally it’s the same.
There is if you've read the code. You can still violate copyright if you're hand-copying an implementation of a feature.
Lawyers recommend "clean room" re-implementations when reproducing copyrighted functionality (when folks bother to ask) precisely to avoid the risk of code being copied verbatim.
Functionality can be copied. Code can not.
Which makes sense when you consider that the sort of code that is getting reproduced verbatim is usually library functions which developers may copy and paste verbatim comments and all into their project, especially when you prompt the AI with the header of a function that has been copied and pasted often, so the weightings will in that instance be heavily skewed towards reproducing that function
> In an attempt to address the issues with open-source licensing, GitHub plans to introduce a new Copilot feature that will “provide a reference for suggestions that resemble public code on GitHub so that you can make a more informed decision about whether and how to use that code,” including “providing attribution where appropriate.” GitHub also has a configurable filter to block suggestions matching public code.
But boilerplate functions don't deserve copyright protection as they are not creative. Can I copyright print('hello world!') if I post it in my repo? Do I deserve a citation from now on?
But there is a desire to understand why an AI provided the output it did (to increase trust in AI generated output), and so there's a lot of study and work going into adding that observability. Once that's in place, it becomes pretty straightforward to identify which inputs to a model provided what outputs.
It's a non-starter for no other reason than potential copyright infringement means the government becomes involved, and they will stomp on the AI mouse with the force of an elephant - the opinions of amateurs and the anti-copyright movement notwithstanding.
As such, AI Observability is a problem that's both under active research, and the basis for B2B companies.
https://censius.ai/wiki/ai-observability
https://towardsdatascience.com/what-is-ml-observability-29e8...
Have a great new years!
Given a black box you can do two things: watch the black box for a while to see what it does, or take it apart to see how it works.
Observability is the former. Useful in many cases, just not here.
If you want to know what LLMs are actually doing, you’ll need the latter. Looking at weight activations for example, although with billions of parameters that’s infeasible.
But irrespective of legality, my point is that attribution is not such a hard problem as people believe, because their mental model of what's going on is much more complicated than reality. I feel they are seeing it as a true AI (mostly non-deterministic and thus hard to observe) and not ML which is deterministic.
One dumb way to do this would be to include self-attributions directly in the stream of training data. So in the cases where the best continuation is to just transcribe the training data, the attributions is included in the data itself.
I don’t think people are even clear on the problem statement here. If I fed in three very similar functions from different sources, and I got a fourth, also very similar, output, what is the “traced attribution” supposed to be?
The case of genuine synthesis is the best case for code generators and aren't the cases where people are expecting attribution.
The problem here happens when the same source is sampled for many tokens in a row because it’s the only match for the context. It could also happen that many tokens in a row each have a long list of sources, but when put together have a subset of sources that appear in every token’s list. That means that someone’s input is being repeated verbatim, even if the network wasn’t trying to reproduce a single source. We could prune the list of attribution sources at the expense of compute by running largest common subset algorithms, which might be sufficient for attribution tracing?
It feels like this whole question might hinge on Fair Use. The network is copying other people’s code one token at a time. We (society & copyright law) all tend to agree that’s fine when it’s a single token out of context, and we all tend to agree it’s not fine when the whole output program matches any single input source. The question naturally becomes, “where’s the line, how many tokens in a row from a single source should be allowed?”
Some code, like the famous fast inverse square root, is widely shared. Even trivial code is often just copy-pasted from a popular SO answer.
Other code is driven by something like convergent evolution. It ends up similar to other code because of the limits of well known algorithms, language syntax, APIs, boilerplate, common coding styles, etc.
In other words, I doubt most human programmers would pass a "copyright check" against all existing open source code.
It’s hard to discuss unverifiable claims of “most” code. A lot of code, maybe most, or maybe not, is proprietary and kept within corporate walls, so we have no idea how much is or is not copied. SO answers are allowed to be copied, by definition. (And at least with SO answers, the code is publicly accessible, and attribution tracking after the fact is closer to possible, right?)
You do have a good point that attribution for some code could be practically impossible to track. This might get a lot worse if we allow AI to remix and republish it, that could even cause feedback loops if we’re not more careful about tracking attributions.
So the challenge for an index would be finding a rare case of creative, important code in a sea of trivial matches.
Incidentally, answers on SO are licensed cc-by-sa, so it's easy to violate copyright when copying them; but no one seems to care, suggesting that small code snippets are indeed "trivial".
Having fair use to be the pillar that all AI training stand on is going to take a while.
Otherwise, if we are going to let companies ignore copyright and licenses because of their pinky promise that their models won’t be recognizable copies of any single input, then we really do need a legal criteria for how much of a single input is fair game, right?
While the very long list would be impractical (but not infeasible), sorting and filtering the list by how much each source contributes would make it easier to manage. Pruning all the entries that are below a contribution threshold would almost certainly shrink it to a very small fraction of the total number of attributions. We don’t need to list all the attributions that contribute to only one token, all we need to know is which attributions ended up with 5 or 50 or 500 tokens in a row, right? Likely not many.
Here's my take from an audio DSP viewpoint.
A composer listens to 10,000 hours of music. One day she writes an original piece based on the annealed parameters (temporal, spectral, intensity, pitch, sequence...) of a million artists. However it sounds like another specific artist who sues.
It is not a cover, a remix, homage, or even forged in the genre... it's just accidentally a bit too like it. (compare: Banana Splits vs. Bob Marley - and - Huey Lewis vs. Ray Parker Jr. - which completely misses the real impact of "Pop Muzik" by M)
The question is, was she exposed to influencing materials incorporated without intent or was it plagiarism (intent to reproduce a derivative etc) ?
By contrast I can take a piece of music, break it down by analysis into melodies, chords, timings, and use FFT to extract the precise spectrums of instruments, feed those to a resynthesis engine that finds new synthesiser parameters to create an exact sound-alike and then deliberately recreate a piece "In the near style of artist X". With a little musical processing I can change the key, inversions, re-template the rhythm to a new swing... always pushing the derived piece into new territory until eventually it's barely recognisable.
Nonetheless, in the second case I have clearly intended to steal someone's idea and "make it my own" by automatable transformation.
In the former case it seems to be a "genuine labour". (whatever that means in 2023)
The genuine artist intends to make something through intellectual labour.
Maybe that's the real question. I mean, about "labour". If the cost of the labour tends to zero, does it really matter?
The weights map inputs to outputs, not training samples to outputs. They are adjusted during the training phase and then kept fixed (until there's a new iteration of the model; and let's ignore online learning for simplicity's sake now, the major models we're talking about here don't use it afaik). From the perspective of a user who inputs prompts and receives answers, the weights are constant. We could create some score indicating how much each training sample contributed to the weights overall, but then that score would be the same for all prompts.
It's really not clear to me what you mean by 'sampling from training data' for a specific prompt. The best I could imagine would be creating another metric measuring the distance between the prompt and all training samples, and then somehow combining that with the first score to get an overall contribution metric for this prompt. But that would be a) quite a bit of guesswork with a lot of modelling freedom and b) quite different from what you described.
Fundamentally, the presence of a next token in some of the samples in the corpus is just as informative as the absence of that token from other samples. You can’t cite one without the other.
If you wanted to give attribution you’d need to list the entire training corpus.
That always seemed like a problem with open source, you should theoretically be able to copy bits of code from hundreds of projects to make a new one, but keeping track of the licenses makes it too much of a pain. So the closest we really see is people vendorizing libraries.
This is a profoundly important distinction.
Back-tracing data to contributory training examples is a genuine "influenced by" relation. Picking the nearest neighbour to a given result (even if its an exact copy!) cannot say anything useful with respect to origins. And given that there will always be some proximate neighbour, it's really a "misattribution machine".
This is bit like how our broken patent system grants or denies ownership of a design based on similarity to extant art but regardless of actual originality.
It's very hard for people to get away from the idea that GPT is "copying" something, but that's not what it's doing. The reality is, to get the exact artifact which produced the code in question, you need "Call me Ishmael" from Moby Dick just as much as the Linux kernel source.
Not always. Sometimes it just copies code without modification.
It never tells you when it does that, though. So to be on the safe side, better assume that it always does.
The easy thing to do is to compare outputs to inputs. This isn't technically hard (i.e. cosine similarity) but it is computationally difficult (i.e. cosine similarity of output to entire input). This would give us some weightings that show similarity. But this doesn't really tell us attribution, rather more a correlation. These are statistical models so there's reason to believe that this is okay.
Then there's inversion. This is processing data backwards through the model. This has different complexities compared to what type of model you're working with. GANs aren't great, diffusions are okay, normalizing flows are trivial (but good luck generating good images from NFs). If we can invert the model it is much easier to investigate and probe for contribution by looking at the distance of the latent generative variable to the location of the latent trained data. Basically you're looking at how different information is contributing to the overall output. This can also be done at every level in the network. Obviously this gets both technically challenging as well as computationally.
Another method would be using dataset reconstruction (this is outside my wheelhouse fwiw). This is where you try to recover the dataset from the final trained network. This too is complicated but there's plenty of papers showing progress in this space (lots of interest from privacy groups).
(TLDR-ish) There's other methods too. But basically what is being said is that there are ways to denote what and how much the training data contributed to the output of the model wherein we can then measure how similar the output is to the inputs (i.e. copying).
As an example: if I take a piece of code you wrote and substitute all of the identifiers that does not create an original work.
But copyright is an entirely different kettle of fish and it is very well possible that Microsoft/GitHub will be able to do an end run around copyright but I don't think they've really thought through what the consequences of that will be. Essentially they are saying 'if you have access to a bunch of source code you are free to slap new terms and conditions on the hosting of that code and then you can use it to train your models, even without active consent of the copyright holders'. In that world Microsoft stands a lot to lose. Much more than your average GitHub repository owner. So this was a terrible mistake on their part.
It actually takes quite a bit of effort to store, then distribute information precisely and broadly. There's a lot of infrastructure, effort, and money involved, and still information degrades and disappears over time.
Any libre information exists because people have put effort into it. Sometimes a lot of effort.
It’s why you can easily find literature written hundreds or thousands of years ago in multiple languages with little effort, but some memo stored in a tape archive from the 1970s is likely gone forever. “Bitrot” is more likely to be a function of diminishing interest rather than increasing entropy.
Little effort, because someone else put in the effort to keep and disseminate it. And we only know of the surviving material, we have no idea how much was lost.
Take the bible for example. It's a classic example of stories from millennia ago making it to the modern world. But its journey took the concerted effort of millions of scribes, translators, and so forth to make the journey. And it has not done so intact. We can even see how the Bible's been changed by comparing just the last two major revisions of it: NIV and KJ. Language aside, there are some pretty major ideological changes based on interpretation. And thanks to some miraculous archeological finds in the dead sea scrolls, we can see how the bible has drifted even further from its original roots.
That's not information being free, it's information degrading and being re-purposed - often re-created - for specific ideologies and beliefs.
I think that indeed, it is the inherent nature of information to radiate itself, i.e. to share (to shine, to spread)
That’s correct. It’s also not what happened here. I don’t violate your copyright when I scrape your web page. I don’t violate your copyright when I train a model on a web page that I scraped.
I violate your copyright only if I use that tool to produce code that violates your copyright. This is similar to how I don’t commit a copyright violation when I read your code, I commit the violation when I produce and publish code that violates your copyright.
> In my opinion it was covered under the GitHub terms of service and is clearly transformative
Now you change the argument completely and let's just say that even with that change that is not how I understand that copyright works and leave it at that. You're very welcome to your own interpretation. Best of luck.
All copyright is a hack; it's not an ideological, internally consistent framework. It is a set of fallible legal rules we invented so that people would get paid for creating things we care about. That doesn't mean that comparisons aren't useful or that we can't extrapolate from the existing rules, but even where fair use is concerned, the actual justification isn't a logical one, it's: "if we didn't have this standard nobody would make anything." So the boundaries around fair use and what things fall into it are not being created purely based on logic or first principles.
The reason why we have copyright is because we want people to be paid for making creative work. The reason we have fair use and a standard of "originality" that treats coders/artists learning from other artists as acceptable is because if we didn't, the entire system would fall apart and nobody would be able to make new creative works.
Everything in copyright exists purely to get people to make more stuff in a sustainable way. It's an outcome-driven process.
----
More recently, there are a lot of people who argue that IP is a real, fundamental property right, but frankly, IP doesn't stand up at all if you think about it too hard. The justifications for why IP theft is theft can't be consistently generalized in a way that applies outside of the IP space. The standards for what does and doesn't count as creative aren't really consistent or based on a straightforward definition of creativity.
A lot of people would love to say that IP rights are just property rights, but... it's not all that convincing, and the history of copyright doesn't really indicate to me that the people building the laws thought of them that way.
And again, that doesn't mean that there's no consistency in copyright rules or that copyright rulings don't have implications beyond the original rulings. But it is almost always easier to think about copyright and almost always easier to understand why copyright laws are the way that they are if you approach copyright as a means to an end, and understand the existing laws not as an attempt to create an internally consistent system, but as a series of attempts throughout US history to achieve a consistent publicly beneficial outcome.
----
With that in mind, I suspect whether or not AI works count as remixing is largely going to be decided based on commercial interests, individual judges, and individual juries, possibly with input from US legislature.
"Everybody remixes" historically hasn't been the most useful argument during these debates? So I don't know how it's going to play out this time around. I vaguely suspect it's going to come down to whether or not individual pieces are recognizable? That's how we got wild copyright laws about some individual chord progressions being treated as derivative works in songs.
exact same problem exists with GPT3 and others.
big tech slashing and burning, ruthlessly exploiting the least empowered people in the tech economy.
neat hack.
Being able to ask a model for help writing Spark, SQLAlchemy, Tokio, whatever, actually increases the usability of free code vs proprietary code.
If an AI being inspired by GPL code that happened to be in its training data because someone other than the creator stuck it in a repo on GitHub is just fine, and if some code produced as a result is practically identical that is fine too, and the resulting code is not considered GPL any more, then the GPL and licences like it are worthless.
The only two designations that mean anything at that point are public domain and commercial. “free” as in “Free” rather than public domain means the same as public domain.
So there is no point releasing free (other than public domain) code, discouraging the act. If I want to control how my code is used at all protected commercial release becomes the only option.
This of course suits the commercial interest behind copilot just fine and dandy…
---
But, if commercial code ended up in the training set the same should apply because in terms of giving the right to use code licences like the GPL and commercial licences are no different: the licence gives the right to use the code. If passing it through an AI gives that right, bypassing the licence, for one case than it should for the other too. I wonder if MS would be happy for copilot's own code to be in the training set and for me to produce and sell something based on the output of an AI trained with their code?!
---
I think the AI systems like copilot should be considered the same as us wet-ware naturally formed intelligence systems in that respect: if it produces something based on code under a particular licence then that something should be subject to the terms of the licence. Ignorance of the licence is no excuse. If the AI can not be made aware of the correct licence and attribution for the code, so it can include that with its suggestions based on it, then that code should not be in the training set.
For decades MS complained about open source code potentially creating this very situation, just with only non-artificial intelligences in the mix, now they are hoping no one can call them on that because it is convenient for them to ignore the issue.
I find most things in all code to be generic and obvious. The same functions are used by lots of other software with just different names that there's nothing novel.
After you've learned to classify lots of functions with big O notation, you will end up writing things in an obvious way for a base-level optimization that fits naturally.
The prospect of knowledge work being resold in this way feels icky. Perhaps I'd feel different if there was a simple, well-defined way to opt out of contributing to the training corpus. Opt-in would be even better.
There are no doubt grey areas and more serious cases as the technology improves and the generated content increases in length and functional value, but I hope we don't throw the productive baby out with the Luddite bathwater...
My personal feeling is that we've gone too far down the road of assigning IP to code and we should be rolling it back. I'd hate for the current AI boom to trigger an extension of current IP law.
If you tell Github Copilot to write a fast inverse square sum algorithm it's allowed to reproduce the idea of the famous quake3 algorithm, but if it produces the same lines and comments as those in the quake 3 source code then that's a pretty open and shut copyright violation
I believe it's "only if they are expressive" i.e. code that is pretty much what anyone would have written in that situation isn't copyrightable.
I think this is the part where we should be much stricter in applying the principle. Similar to "novel and non-obvious" in patent law, it's become diluted in favour of IP overreach.
It's not at all obvious to me why the training set copyright wouldn't nominally flow through the model, even if that seems impossible to actually implement in practice.
ChatGPT is just a really impressive word-blender. All of those "sufficiently distinct" snippets are in a meaningful sense recombinations of other copyrighted snippets it was trained on. When you're in a very sparse part of the space or have a single work that's present many times in training like fast inverse square root, the model may even replicate recognizable chunks of the training set. It could be that this is sufficiently transformative that the output is not derivative, but as I said it's not obvious to me.
This might be the worst metaphor I've stumbled across in this whole debate. You're surely making an argument here for the opposite team.
For what it’s worth, I wish the approach was as you describe but in all cases: make it impossible to defend attribution of basic building blocks regardless of how deep your pockets are.
I had a “DevOps” project I was working on creating deployment process using AWS technologies (disclaimer: where I work in Professional Services). I needed a few relatively simple Python scripts.
I first asked ChatGPT:
“given a JSON file like this [{“company”:”${company}”} replace the word surrounded by ${} with the equivalent environment variables using Python”.
It worked perfectly. But it hardcoded the input file.
“Modify the script to accept the input file using a command line argument -json-file using argparse”
That worked and used the “required=true” parameter.
“Instead of skipping replacement if an environment variable is not found, raise an exception”.
And it seems to understand the AWS SDK. I told it to convert a script I wrote to a CloudFormation custom resource using the cfnresponse module and it knew the correct Lambda event structure, the event format etc.
I have seen reports and have witnessed it making up functions occasionally.
But, I believe with the right prompts, it should be able to create simple CRUD scripts.
I know we’re “just” talking about code here but decisions will be far-reaching and if we let the powers that be force AI-generated content to be copyright-attributed, two things are going to happen:
1. The biggest benefits of AI are going to be pushed back for decades and possibly indefinitely
2. Artists will be absolutely slaughtered as the same rules will come at them full-force. Almost every artist draws inspiration from other creative work. That’s how creativity works. Can you imagine how stifling it would be for every artist to have to document and “pay for” every piece of art they’ve ever seen??
https://news.ycombinator.com/item?id=34242124
If AI can produce patentable and copyrightable proteins, the first pharmaceutical company to press a button will own all possible proteins. And then all possible permutations of every undiscovered drug.
If we let the powers that be force AI-generated content to be copyrighted, two things are going to happen:
1. The biggest benefits of AI are going to be pushed back for decades and possibly indefinitely
2. Artists will be absolutely slaughtered as the same rules will come at them full-force. Almost every artist draws inspiration from other creative work. That’s how creativity works. Can you imagine how stifling it would be for every artist to have to document and “pay for” every piece of art they’ve ever seen??
But actually carrying attribution forward is going to be hard. Things that come out of a generative model that are substantially identical to a particular input do so for one of two reasons: one particular training input to the model is overwhelmingly the "best" source of response to the prompt, or a training input has been repeated over and over again in the training set so as to become the consensus response to one or more prompts. The first might reasonably feasible to track down, although it's bound to be computationally expensive. The second ... really tough, since it forces the model to "know" which of those many training sources is the one that should attributed. The widely circulated example of an image generation model reproducing Steve McMurray's famous photo of Sharbat Gula, a green-eyed Afghani girl that appeared on National Geographics cover in 1985 shows the problem. Do a Google search for "green-eyed Afghan girl" and you'll find hundreds of copies of varying resolution and definition, and hundreds more derivative versions, of McMurray's photo. A model spitting out yet another derived, but nearly identical version is likely drawing from hundreds of those images itself, not some original root, copyrighted golden copy. Which should it attribute?
I know that even non-copyleft licenses like the Apache and MIT licenses are copyrighted and require attribution however it would have caused far less controversy than training on GPL licensed code.
Now with all that, there really isn’t anything here to get worked up over.
Is this more than a opinion?
Because you can have whole programs as a long functional expression (not that I am a fan of such a coding style, but it exists).
For example, if I asked you to write Java code to print the sum of two integer variables named a and b, followed by a newline, you'd probably write pretty much exactly this:
System.out.println(a + b);
You might vary the spacing around the punctuation, or a few other unimportant details, but there's not much room for creativity there while still accomplishing the desired function. It isn't copyrightable.
And this is a good example of how the word functional means something non-mathematical here: that isn't even a functional expression in the mathematical sense, since there's no return value and the printing is a side effect. It's purely imperative.
Opinion me, this statement could apply to the vast majority of copyright law in the US.
It's also going to be hotly debated in the future, because right now most of the commercial AI-generators are just kind of ignoring this and at some point I think they're going to make the argument, "this law/interpretation can't hold or else it would be commercially devastating for us, and it needs to change."
But under current US copyright law, the US copyright office has pretty consistently ruled that:
- purely functional expressions without a creative aspect can't be copyrighted (ie, you patent an invention, you don't copyright it, and stuff like recipes aren't eligible for copyright at all, only the surrounding text is).
- AI-generated content does not have a creative aspect and isn't eligible for copyright protection.
They've even gone so far as to revoke copyright protections to AI-generated content[0][1]. We haven't gotten a similar ruling about code that I'm aware of, but given that the US copyright office already only grants copyright to code under the assumption that coding is a creative act, it's very difficult to imagine code having more legal protections than an image or a book.
----
In my opinion, this is the much more interesting debate than model's source[2], because nobody is really talking about it, and depending on how the inevitable legal challenges play out, it could completely change the field. It's an interesting choice to think about -- let's say an AI gets good enough that you can sit down, describe an app, and the AI just generates it straight up. Would you still use that tool even if the resulting code was public domain and you couldn't assert ownership over it?
Which would be more important: the accessibility/ease-of-creation, or owning it?
It's also interesting because I haven't heard a ton of good legal arguments from people about why this shouldn't be the case, so there's a nontrivial chance that it holds in the future. I've heard tons of logic-based arguments: "it's creative because I decided what to generate, photographs are copyrightable, etc..." And those aren't necessarily bad arguments, they're just arguments that the US copyright office has already rejected. I haven't seen a lot of arguments about "here's why AI generated code would be eligible without a legal challenge overturning existing copyright rules."
----
[0]: https://www.theverge.com/2022/2/21/22944335/us-copyright-off...
[1]: https://aibusiness.com/ml/ai-generated-comic-book-loses-copy...
[2]: Not that the debate over the training data isn't interesting or important; that could also have implications for stuff like fair-use in transformative contexts. It's just that I think it's less likely to have far-reaching implications or be upheld.
That's overgeneralisation. A language model alone, yes, is just derivative. But a language model trained on solving problems with reinforcement learning can surpass humans. For example AlphaGo and AlphaTensor are models that learned from running simulations.
Under current copyright law, if you built a model that surpassed humans and trained itself on simulations and it wasn't making derivative works but instead producing purely original output, the resulting output would not be copyrightable.
In fact, under current copyright law, if you built an AI that was sapient, the creative work it made would not be copyrightable -- unless you literally got the law to grant it legal personhood or something.
Moving AI in a direction of being more fundamentally creative makes this harder, not easier. The legal defense of copyright in cases involving AI is that the creativity is coming from the human using the AI, not the AI itself. As the AI gets better at generating usable output with less directing and micromanaging, that argument gets weaker, not stronger. The question isn't whether or not the AI is creative or produces novel output, the question is whether a human exhibited enough creativity making the work for it to meet the standard of gaining copyright protections -- which, traditionally, the US copyright office has usually held isn't true for AI-generated content.
AI needs to be regulated somehow. Precedent needs to be set.
I struggle with this one. How are these models different from the typical human creative process of:
1. look at lots of existing art to get inspiration
2. select components from several different styles & add your own flair
3. call the output an "original" painting in your own style
Of course I see the other side as well. These models are just a large multi-dimensional function of all of their input data, which means they ought to be derivative works of all of their inputs, and their outputs should thus also be derivative works of the inputs.
It feels unsatisfying (and a bit naive) to say that filtering a bunch of data through a neural network somehow clears the copyright of that original data, but on the other hand, the AI revolution depends on that interpretation.
But they directly reproduce the source material.
AI art they clearly does not.
The luddites seem a better parallel when it comes to scale. Where a machine comes along capable of producing in much higher quantities and in much greater efficiencies.
Or perhaps photography? Also a fear of scale. For a long time photographers were not considered artists. And there were calls to tamp down on this innovation for the harm it would cause.
I don't suppose anyone back then could imagine now how photography is used. Perhaps they would feel it much better that a tiny artist quickly sketch out the thing my phone was looking at.
It does not 1-1 recreate source material, but if we didn't have human artists anymore and replaced them all with the current stock of AI we would have no more developing art movements. It is wholly uncreative in a way that humans are not, and it does not understand what humans would find interesting, only what humans have already made. This makes it a pretty good tool for some amateurish applications or places where quality isn't important, but it is not a 1-1 copy of a human capable of making human art.
The work I've seen manipulating, weilding, and modifying this new tool, has been very creative In painting, reprompting, texture and uv map generation, dreambooth and dedicated fine tuning. Integrating these workflows into human directed procedures to produce in hours what would take days, sometimes months, sometimes not even within the scope of human capability.
If you think the only thing here is people jibbering at a black box and posting 512*512 images, I can see why dumping the entire thing would seem reasonable.
'but how is capturing an image with a camera lens any different from capturing it with my retina? both are simply methods of refracting light and translating that information into a storage medium. the camera is just remembering* like a human remembers and we mustn't let luddites regulate this revolutionary technology or we'll impede the progress of the camera becoming an even more better rememberer'*
the social ramifications of the camera are still being reckoned with to this day and yes, cameras very highly regulated compared to pretty much any other human method of expression or transcription, in fact we are still coming up with new things you cannot film eg the recent crop of revenge porn laws.
Finally, someone has put into words the logical end of the trite photography analogy.
It's a refreshingly sane perspective to see in this thread.
https://twitter.com/kortizart/status/1588915427018559490
I would think these are close enough that any human that produced that output could be claimed to have plagiarized.
Famous images prevelant in the dataset that are trained against many times and the directly requested are produced as is, because that's litterally what was requested and the model was trained against them.
There's only one way to draw The Mona Lisa. These AI's are not used for this purpose, nor is any of this core to their functioning. And I have to assume you know that.
> to shrug while others lose their livelihood and have their self expression systematically duplicated and commodified.
The same exact arguments were made when the camera was invented.
In order to be consistent, you would have to argue that the camera should have been banned to protect painters.
> That you can't empathise or find it 'naive'
Oh we understand. We just know that it is the exact same argument, that ludites make every single time a new technology is invented, that increases efficiency.
It is always the same argument repeated over and over again. And if those arguments weren't good in the past, we aren't going to listen to defeated arguments now.
There are two kinds of data - the idea and the expression. You can protect expression, but can't stop AI from learning the ideas. This is not "clearing the copyright", it is fair use of the data. Eventually it will even learn how many fingers and heads to draw on a human - the kind of "idea data" you can't copyright.
We should stop mixing together idea and expression. Ideas are not owned by anyone. The real question is what criteria and threshold to use for judging infringement, so AI people can get on to making safe models.
Besides ideas, you can't copyright recipes, APIs and purely functional/obvious code.
swiped from unassuming creators who put it online for everyone and Google Bot to see? I presume they follow robots.txt when they crawl
And maybe some of them even read HN occasionally and are currently enjoying the article.
And the more meta question is whether the creators' rights should even extend to that realm. Can you specify in a license "This text must not be read by students in the context of an educational course" (which is arguably the closest analogue to AI training)? If that's not possible, then where do we draw the line between tools-assisted learning by humans and human tools learning on their own? You can arrive at answers, but it seems to me like there are many possible reasonable interpretations.
It would be more accurate to compare those students with the corporate interests involved in AI training, not the model's use of the content.
Can we all legally pirate educational books since it's for self training and producing transformative outputs? Can I consume all media (books, movies, music) the same way, and call it fair use? Also for software?
The debate is about content that is generally legally obtained, but might come with certain restrictions. Where restrictions is a broad term and might also just come in the form of a copyleft license, eg. My main point was that in many situations, eg involving open source licenses, it's really not clear from the terms what the creator's intent regarding AI training was. And the broader question is whether training is fair use, or something else, maybe even a new legal concept that would have to be established. Or, what's the difference between art students going to the museum to be inspired, and Dall-E 'looking' at public domain images?
Regarding your point, I wonder how different is it to break the "terms of use" for fair right, vs breaking the "terms of obtaining" alltogether. They both are about jumping over owner's will, who has the full rights of the work.
I very much look forward to see how the ethics and law around these issues evolve.
Otherwise any deluded fool with extreme views on copyright could claim ownership way beyond that which the law currently offers.
If only we held all myths to such high esteem as we societally hold ownership.
but to do that without any reasonable social support nets is absurd
if we take that
> The only way to truly own something, is to either share it or destroy it.
could I argue that then, we would be taking ownership of the concept of ownership??
ahahha... I think this is kind of funny. And I'd admit that it's not very helpful to the goal of revising the very concept of ownership.
my own parent comment is already controversial with many replies and exactly 0 points (at the time of this response)
You don't mind if I borrow your car do you?
And your house?
Let's not quibble about the particulars of me ever giving them back.
GP did not limit the scope of their opinion on ownership to IP.
As I understand so far the main reason to seek a revision of the concept of ownership is exactly due to the existence (enabled by internet technology) of digital assets.
copy-pasting is HOW computers work. copy-pasting does not do well in society ruled by the exclusivity-mindset inherent to marketplaces (and their societies) of tangible assets
I agree with you that digital "asset" ownership is new, unexplored territory for humans.
We've never been able to separate the content from the distribution medium before now, and we're struggling to recreate a model we're familiar with (physical media) by imposing absurdities like DRM.
NFTs are also an absurd way of trying to solve the same problem.
The issue we wrestle with is mistaking "ownership" with a access.
I can own a physical book, but I do not own the content of the book, I've paid for access to that content and the book is the medium.
Digital assets are the same, I don't own them, they are a form of access to content that someone else owns, and has granted me an either implicit or explicit license to use.
I think this is the concept of ownership that needs to be revised.
Rather than a mp3 being treated, and thought of, like a book - it should be thought of more as a movie ticket. Something that grants me access to content within the limits defined by the content owner.
I'm trying to point out that this distinction (ownership as distinct from access) leads towards a capture of the digital advantage by people with better leverage.
moreover, I disagree that if I own a book I do not own the contents of it. the mindset that I don't seems too close to saying that I can know things (well understood learned concepts) but still somehow not own them.
This in my view is like a 'hook' which pulls towards the reality that somebody else owns the contents of my own mind, hence that I do not own that part of myself. I hope you see where I'm going with this and why I find your posture troubling.
It's only a few short (conceptual) steps from doing away with individual freedom for the sake of what?
ZipCar, et. al.
> And your house?
AirBnB
> Let's not quibble about the particulars of me ever giving them back.
Symmetry.
Granting overlords the power of ownership, reducing individuals to renters that must comply with the whims of the property masters, is not a path to increased liberty.
Sure. I mean, I'm not an economist nor a sociologist, I bet there's some nuance to that, eh?
> Granting overlords the power of ownership, reducing individuals to renters that must comply with the whims of the property masters
Peasants and serfs. In re: technology I talk about Morlocks and Eloi. Yeah I think that's a lousy way to structure society.
The original point that you replied to was, "ownership is a concept in dire need of revision" and it sounds to me like that's where you're coming from too?
Everyone must unlearn the term "Intellectual Property". These laws are anti-property rights. They are Intellectual Slavery laws (https://breckyunits.com/an-unpopular-phrase.html).
The United States government employs more knowledge workers than all other companies (see NIH, DoD, CDC, NASA, NOAA, NWS, et cetera). Everything they produce is public domain, by law. And yet, the people producing these information products still get paid!
We don't need (c)opywrong laws. We don't need Intellectual Slavery laws. We still have cotton even after the 13th Amendment (we actually have more and better cotton now), and we will still have creative works after the passing of the Intellectual Freedom Amendment (we actually will have more and better creative works) - https://breckyunits.com/the-intellectual-freedom-amendment.h....
I suspect that the folks who invest hundreds of millions of dollars into production costs for a movie, rather enjoy the ability to recoup those costs by restricting access to only those who are willing to pay for the privilege.
do they have a right to enjoy being rewarded for work they didn't do but their ancestors did?
if the intention of those laws was to encourage creativity, how come they're stiffing it more than encouraging without any signs of any government trying to correct the laws to better match their purported intentions?
We should have stuck with the original 14 + 14 as laid out in the constitution.
Did you not read: "The United States government employs more knowledge workers than all other companies (see NIH, DoD, CDC, NASA, NOAA, NWS, et cetera). Everything they produce is public domain, by law. And yet, the people producing these information products still get paid!"
These laws are counterproductive and have terrible second order effects, and are cruel to everyone except a tiny % of the population, just like slavery laws.