Please don’t upload my code on GitHub
nogithub.codeberg.page
nogithub.codeberg.page
Fundamentally it seems to me that Copilot truly does synthesise code from some store of knowledge (even if it's hard to understand what this store of knowledge is), and the problem is that it's synthesising code identical to existing code. There are legal tools and also rhetoric that are designed for dealing with this problem of "synthesising something that is identical to an existing thing", and it's different tools and rhetoric from the ones we have for dealing with the problem of "stealing or copying existing things". It's valid to have an issue with with Copilot ingesting your code, but unfortunately people are largely using tools and rhetoric from the latter category to approach the issue, and that slight misapplication is causing their issues to fall on deaf ears a lot of the time.
I expect that the copyright holders could reach to me, tell me that my code is identical to theirs, believe or not that I independently created the same code, and at least ask me to stop distributing it.
Of course
t += total(e)
is easier to defend than monthlyTotal += dailyTotal(expenses)
because the chances that I didn't copy their code get slimmer and slimmer as the names of the identifiers and the structure of the code get more complex.If I actually looked at their code, it would be wiser from me to at least change the names, maybe also a little bit the structure of the code. That's the difference between being inspired by something and copying it.
TL;DR: if Copilot generates a copy, it's a copy.
Then you're in the clear, though you may need to convince a court that that's what happened. (Patents and trademarks could still be issues, but there's no copyright issue)
It's a common mistake because we are not used to LLMs.
If the LLM reproduces a significant portion of a program token-by-token, is it a derivative work and is it not fair use?
A derivative work is not fair use. If you end up with a significant portion of another person's program in yours, (such that a substantial portion of your program is in some way related to their program), that will likely be a derivative work - but the definition of a derivative work depends on your jurisdiction and use-case. If you're unlawfully using the source material to produce a derivative work, you cannot copyright that derivative work under 17 U.S.C. § 103(a), and under the same section, you can only copyright your modifications, not the original.
It would be hard to argue fair use in this case; fair use only really applies for parodies/criticism, reporting, and scholarly works - and generally that's an affirmative defense, rather than an express or implied right you have.
Honestly, Copilot is difficult because Copilot can't be the author of the code; the person who used Copilot is the "author" of the code, and I think they'd be the ones liable for copyright infringement if copyrighted content ends up in their code.
To argue someone performed copyright infringement, all you need is to prove (1) a valid copyright exists; (2) that the person had access to the work; (3) the person had the opportunity to steal the work; and (4) that protected elements of the work had been copied (afaik generally under a "substantially similar" standard). Copilot offers an easy way to check both (2) and (3) - a copyright holder could argue that people had access to their code through Copilot, and that Copilot offered an opportunity to steal the work.
I think of it the other way around: I'd be more trusting that LLMs aren't a problem in that sense, if the big commercial entities linked to them used their own code to train them as well as public repositories.
MS's codebase must contain much good training material, unless they think their own code is crap and are embarrassed to let the AI look at it!
I am still undecided, but erring in the direction of not wanting my stuff to be used to train them, though I've long since passed Douglas Adam's "35 year's old" tech barrier so maybe I'm just not liking change…
For now at least I won't be putting my own code in services like GitHub, then at least if someone else does or an LLM is trained on public sites generally, I've not explicitly agreed to anything that says the company can use my code that way - which you do (they say, I believe there is at least one court case brewing on the matter) when you sign up to the services (agreeing to their terms in the process) and use them that way.
Having worked at Microsoft, I'm not sure I want my coding assistant trained on Microsoft's codebase.
> legal tools and also rhetoric that are designed for dealing with this problem of "synthesising something that is identical to an existing thing"
I feel like I’m reading the exact same sentence that someone posted about if StableDiffusion included images or was just “learning patterns” and totally generative from 6 to 8 months ago.
People argued, oh so hard about it.
“But stable diffusion is GB in size, how can it have embedded full images of a dataset that is so much larger? It’s not possible.”
…but, it turns out it is possible (1) and it does, in fact, have embedded full copies of the source images.
Now. Let’s talk about code…
> and the problem is that it's synthesising code identical to existing code
There’s a word for that. It’s called copying.
Copying fragments, copying full text. It’s a black box, it spits out content that is indistinguishable (again, see stable diffusion, a lossy jpg of an image is the same image even when it is not bit-for-bit identical) from the input.
That’s copying.
The black box might be a neutral network, or a python script that reads a file from disk. The process is irrelevant.
If you memorise a block of code and type it out by hand with no reference, you are copying it.
These models copy code.
These models have embedded full text copies of some code.
What we do with that is an ethical question, but that it is true is not in dispute. It is true. It’s documented.
If you think these models are not copying at least some training data as output you are factually, and provably incorrect.
(1) - https://arstechnica.com/information-technology/2023/02/resea...
Black box. Input. Output. Is the output the same as some training input?
It’s not rocket science.
“…but humans…” argument is not relevant. This is not a human.
It doesn’t matter if there is a human doing it or not.
For your supposition to work, the input to the box would be only an abstract summary of the logical steps, and the output an exact copy of some other thing that was never an input.
In that case, yes, it would not be copying.
..but, is that the case? Hm? Is that what we care about? Is it possible to randomly generate the exact sequence of text with no matching training input? Along with the comments?
It seems fabulously unlikely, the point of being totally absurd.
Machines don't. It doesn't matter how fancy they are. The law doesn't care.
So yes, a human in a black box is different from a machine in a black box until laws change.
Wrong. Literally the whole thing about law - and especially about intellectual property laws - is that process is as much, if not more relevant, than the outcome. This is why "code as law" efforts are plain suicidal. This is why you can't just print out a hex dump of a pirated MP3 file and claim it's not copyright violation because it's just a long number that your RNG spit out - it would've been a good argument if your RNG actually did that, but it didn't, and that's what matters.
This is what it means when we say that, for just about everyone except computer scientists, bits have colour[0]. Lawyers and regulators - and even ordinary people - track provenance of things; how something came to be matters as much, and often more, than what the thing itself is.
This is what makes generative AI case interesting. They're right there in the middle between two extremes: machines xeroxing their inputs, and humans engaging in creative work. We call the former "copying", and the latter "innovation" or "invention" or "creativity". The two have vastly different legal implications. Generative AI is forcing us to formalize what actually makes them different, as until now, we didn't have a clear answer, because we didn't need one.
--
The court did not agree. They looked instead towards an anti-biker gang law that illustrated that a biker bar can be found guilty of assisting with gang crime, even if no specific crime can be directly associated with the bar.
The defense team argument - that prosecutors need to prove that a crime had occurred - failed. The courts only require that the opposite is not believable, which given all the facts around the case was deemed sufficient. In that question the process doesn't matter. If the court do not think it believable that copying has not occurred, any argument about "machines xeroxing their inputs and humans engaging in creative work" will be ignored.
They don't have the same rights to register a copyright.
Are you trying to make money off of it? Is it for personal use? Was it protected or public domain?
It just all depends on the circumstances. Human or not.
If you cover a song on YouTube, you can apparently be demonetised or attract copyright strikes from the original artist. But in software, you tell me an algorithm and I code it up, that’s an original work.
The lines all seem pretty arbitrary to me.
https://www.nationalgeographic.com/science/article/autism-ar...
Stephen Wiltshire is cool.
Any company that is using copilot or similar is walking right in to a massive legal minefield. Any engineer that chooses to use Copilot without explicit approval from the company should seriously worry about what target they're putting on their back. You don't want to be putting yourself in a position where your employer can say "They did this without our approval" or worse "Despite explicitly having told them not to", should there end up being legal action. Someone is going to establish precedence one way or another at some point, don't let it be you.
You mean, like the example the article used?
https://web.archive.org/web/20221017081115/https://nitter.ne...
"sparse matrix transpose, cs_"
Generated the authors matrixes code.
https://codeium.com/blog/copilot-trains-on-gpl-codeium-does-...
It produces an almost exact copy of some training data.
> These models copy code.
These assertions have very little weight - what matters is what the law says. And for now, in the jurisdictions that matter, the law is still silent. Maybe it will catch up via either legislation or case law, or maybe it won't. But for now there is no legal consensus on what these new things are doing.
Of course there is a moral/ethical component to this, and individuals are obviously free to hold personal views. But unless some global working consensus forms (unlikely) then I don't see how change will come from that.
I suspect that it's going to be difficult for politicians to find the will to do something about this issue, and then to get their heads around it enough to make good laws.
The legality aspect that you are injecting into the discussion is irrelevant
This is possible. For example, many functions f(0) return 0; if we can derive a function g(x) that generates y without access to T (ie. the training data) then we can plausibly assume that f(x, T) might indeed not be copying from T.
…but can we do so?
If so, we must admit that it is possible that f(x, T) = g(x) and no copying is taking place.
So, in the past 50 years of prior work can we come up with an example of a function g(x) that generates the exact code we see coming out of these LLMs?
I’ll grant, it’s not impossible.
For example, if stable diffusion generated sin wave patterns or fractals that exhibited “deep complex structure” and it happens that some artist had written code to do that exact thing, there would (I believe) be a fair case that stable diffusion was not copying the artist, it was simply a parallel implementation of the same generative code.
However, now my scepticism kicks in.
For code written by hand, that was not generated, we are suggesting that a LLM is a parallel equivalent implementation of a human mind, and it can, with purely “learnt patterns” replicate not only the intent and structure of known code, but the exact code itself, repeatedly.
I’ll grant. It’s not impossible, and it’s very difficult to prove, but it seems fabulously, unbelievably, extraordinarily unlikely that a clean room reimplementation would be identical, repeatedly.
We’re entering into the domain of assigning probability and then doing a zero knowledge proof here, but the solution is easy.
Just prove by counter example.
Train a model that generates such code without the code as training data.
You have now won the argument.
…on the bright side, you’ve also solved the 10 million dollar question of “how do I train my LLM when I don’t have enough training data” (because that’s what you just did; train it to do something without training data), so you’re now rich and give zero ducks what I think. Congrats.
> For example, if stable diffusion generated sin wave patterns or fractals that exhibited “deep complex structure” and it happens that some artist had written code to do that exact thing, there would (I believe) be a fair case that stable diffusion was not copying the artist, it was simply a parallel implementation of the same generative code.
I understand your pseudo equations and I understand [some] engineering.
Why are artists writing code? Is that to show the parallel of if artists did write code?
Unless artists do really write code but that’s beyond the scope of what I’m asking.
And that in that case, g(x) that can produce a perfect copy of the artist's work without having the image in the training set provably exists, and is the actual mathematical fractal function itself.
Basically if you take something then feed new info into it is that new thing a new thing or a derivative of that work or other? With these models the data and mathmatical spline is so smeared across thousands of endpoints it is hard to tell. However, I do not think the courts have to think about the details of how the models work. They can do something else so a jury can get its head around it.
But basically the courts will probably simplify it into copyright things go in one side. Other things come out the other that sort of kind of resemble the original thing. Do not worry about the details of how it is done. Is that new thing owned by the original copyright holder? Or many holders, as it is smeared into a hundred other items that may or may not have copyright holders? Or is it a new work? Or is it a derivative work just colored?
This could go either way. As someone put it very nicely here a few weeks ago. These AIs are like the most amazing auto complete you have ever seen. Now in copyright if you make something and I independently come up with the same exact thing. Also I can prove that I am in very little trouble. Now in this case without that input code that AI model probably would not predict that exact string. But it in some cases exactly predict it again but can predict thousands of other things also. Is that prediction copyrightable? What if it predicts part of the code but with something else? I as a 3rd party who did not create the model and just used it and got the code what are my liabilities? Then if it is what are the consequences of that? The courts will have to decide eventually.
I believe China is trying to limit AI powers but who knows how that will go.
> …but, it turns out it is possible (1) and it does, in fact, have embedded full copies of the source images.
The _not possible_ assertion is in response to people arguing that generative image models literally copy and thus "steal" the images in their training set or that such generative models "work" by infringing on said images. The assertion is not that the models cannot possibly contain _any_ copies of training data, the assertion is that the models are generalizing and can not contain a majority of the training data.
Now, let's look at some lines from your linked article titled "Paper: Stable Diffusion 'memorizes' some images, sparking privacy concerns", which happens to lead with scare quotes around "memorizes" and already hedges with "some":
> Researchers only extracted 94 direct matches and 109 perceptual near-matches out of 350,000 high-probability-of-memorization images they tested . . . resulting in a roughly 0.03 percent memorization rate in this particular scenario.
I do not consider a 0.03% rate of full copies particularly damning. Instead, it sounds like a flaw in training which could be addressed, which the article does in fact also mention being as a matter of overfitting to images that are over-represented in the training data.
> the 160 million-image training dataset is many orders of magnitude larger than the 2GB Stable Diffusion AI model. That means any memorization that exists in the model is small, rare, and very difficult to accidentally extract.
Which side's claims do you feel this line supports?
Generative text models should, in theory, also be generalizing and not memorizing. In practice, however, it does appear far easier to get Copilot and the like to spit out exact copies of code, complete with license text. Presumably, there's just not enough code out there to sufficiently generalize from, so we're seeing overfitting as a matter of course.
Or hell, the extracted potentially infringing and over-represented material might all be pop music set to variations of Pachelbel's Canon, and I'd pay to see that lawsuit.
1. The copyright on the composition. This can also include arrangements - for instance, Gershwin's original piano version of Rhapsody in Blue is now public domain, but the orchestral version everyone knows,
2. The copyright on the sheet music (the actual layout, spacing, editorial notes, things like that.. it's actually an insanely deep subject. I've got an 800 page book on the subject - which is referred to as music engraving, as up until about 40 years ago it was literally done by engraving the plates by hand. Much much harder problem than doing normal book-style text layout, as it's fully 2D, whereas text is basically 1d with occasional special cases. (NB: This copyright is really only relevant to the musicians, conductors, etc, but it does matter.
3. The copyright of the particular recording. This is the really relevant one. A 5 year old recording of a 500 year old work is very much under copyright.
No one sued Napster because their guitar tabs were being shared.
From GP: " they may largely consist of performances of public domain compositions."
My entire point is that the composition being PD does not mean RECORDINGS of it are PD.
few images of of millions
If Edison breaks into Tesla’s house at night, takes* the invention off his workbench, goes to his factory and uses it as a reference to build and sell the same thing, then Edison has stolen and copied Tesla’s invention.
In both cases we can say “stolen and copied”, but actual sequence of events and the legal recourse that Tesla has is different in each case.
This difference matters because open source is like Tesla having his workbench in a public place where people are encouraged to watch him work, play around with the inventions, help him build, etc. In the “Tesla’s public workbench” situation, Tesla’s recourse is debatably impacted in the first scenario, but it is definitely unchanged in the second scenario. Arguing as though Copilot is doing “second scenario stealing and copying” rather than “first scenario stealing and copying” sidesteps that debate about whether open sourcing your code made it ‘up for grabs’ for natural and/or artificial intelligences to learn from and/or memorize. Now, it’s not clear what side will win that debate - I personally think “this has infringed intellectual property rights” has a decent chance of being the victor. But sidestepping that debate is setting off peoples’ “you’re trying to pull a fast one” alarm and making them unsympathetic. That’s what I’m getting at here.
*: to avoid “steal vs copy depriving of property”-type discussions, postulate that Edison returns the invention to the workbench before Tesla notices
As to the linked article, I would quote from the article itself: “[The paper] is dense with nuance that could potentially be molded to fit a particular narrative”. I could say more: both the article and the paper state there were zero byte-identical matches, and 203 direct or perceptual near matches out of 175 million generations (or out 160 million training images) is 0.0001%, which argues against this very strong claim of “factually and provably incorrect”. But I think the argument I will actually make is this: I said Copilot was synthesizing identical code from its knowledge base, which is completely compatible with the claim that it is outputting code from memory (in the human sense of memory). What it is not compatible with is the claim that it is outputting code from memory (in the computer sense of memory).
AI instead has looked at all the existing code and it has learnt to do no more than what the existing code can do.
Do you see a difference?
Humans make sense of the world with an extremely limited amount of information. I have not read all of GitHub. Comparatively, I have not read a fraction of it. And I could not read all of GitHub even if I wanted to.
However, I do not need to read all of GitHub anyway because, as a human, I am capable of understanding.
Current generation "AI" is not capable of understanding. Current generation "AI" merely computes the probability of any given word appearing next in this sentence based on an utterly absurd amount of raw data.
That is not what humans do at all. We do not need this information. We cannot even process it.
> Do you see a difference?
Put another way: Some Markov models are Turing complete, so "merely" computing probabilities can be Turing complete with only minor steps, so trying to downplay the potential capabilities of models like this by handwaving about things we don't know is foolish.
We don't need that scale of input, but we also don't know if LLMs need that much input to do well, or if our current training protocols are simply poor. With ongoing work on reducing the training cost, it is at a minimum clear that current training methods are far from optimal.
Are you sure about that?
Yes I'm quite sure. Unless you're claiming that source code existed before humans?
It sounds like what you are saying is that the AI cannot write code which would do something it hasn’t seen in the training set. That is how I interpreted the “it has learnt to do no more than what the existing code can do”. Do i understand you right?
If so, how do you know that? Are you talking about the limitations of a particular implementation or a limitation of all AIs as a concept?
No, the problem is that it's illegal to do so.
Copyright law applies in both cases. If you create work that’s substantially similar to an existing one, you’re risking copyright infringement even if none of the original work’s content was reproduced in the mechanical sense.
If this were not the case, why would cover bands pay the original artists for rights to the songs?
“I learned to play ‘Yesterday’ by heart” doesn’t mean you can do anything with the song without paying the Beatles. The same applies to machine learning, if all the model has learned is to imitate copyrighted works.
If it was possible to point where the pixel values are actually stored in a jpeg file (i.e. run dd on the file and it spits out the pixel values), I would be a lot more sympathetic to the view that LLMs are not stealing code.
If you decode the jpeg file you do get the pixel values.
One could describe llms similarly.
I don't really care whether there's a lossless, lossy, or AI compression of my stuff. I care that it's my stuff.
The only leap here is the fact that the programmer has outsourced the learning to a tool that does it for them, which they then use to do their job, just as before.
Nobody is arguing against uploading code. It's about Github/Microsoft specifically.
However, at a first glance, it still feels to me like an unavoidable reality that if you publish source code it'll eventually be ingested by Copilot or whatever comes next.
I mean, for the rest of the content all the new fancy LLMs have been trained with, there wasn't a Github equivalent. They just used massive scraped dumps of text from wherever they could find them, which most definitely included trillions of lines of very much copyrighted text.
In short: not only I don't really see an issue with Copilot-like AIs learning from publicly available code (as I described in the GP comment) but I also think if you publish code anywhere at all it's inevitable that it'll end up in Copilot, regardless of where you host it. If you want to make it more expensive for Microsoft to scrape it, sure, go ahead, but I don't think it matters in the long run.
I’d be quite careful with of this view.
By your logic, it should be ok to take the Linux kernel, copy it, build it, then sell it and give nothing back to the community that built it. Then just blame it on the authors for uploading it to the internet ?
So that’s not accurate.
Though, GitHub would do well to also bake-in approp attributions if a significant portion of the generated code is a copypasta.
In the case of ChatGPT-x, Open AI company which is disguised as a not for profit with a goal of producing ever more powerful models that may eventually be capable of replacing you at work while seemingly not having any plan to give back to those who’s work was used to make them insane amounts of money.
They haven’t even given back any of their research. So it’s ok to take everyone’s open source work and not give back is it ?
This isn’t some cute little robot who wakes up in the morning and decided it wants to be a coder. This is a multi-national company who has created the narrative you’re repeating. They know exactly what they’re doing.
The interesting debate should be what happens in the gray area, when you read a lot of code and learns patterns and ideas.
> This can lead to some copylefted code being included in proprietary or simply not copylefted projects. And this is a violation of both the license terms and the intellectual proprety of the authors of the original code.
If the author was a human, this would be a clear violation of the licence. The AI case is no different as far as I can tell.
Edit: I'm definitely no expert on copyright law for code but my personal rule is don't include someone's copyrighted code if it can by unambiguously identified as their original work. For very small lines of code, it would be hard to identify any single original author. When it comes to whole functions it gets easier to say "actually this came from this GPL licensed project". Since Copilot can produce whole functions verbatim, this is the basis on which I state that it "would be a clear violation" of the licence. If Copilot chooses to be less concerned about violating the law than I am then that's a problem. But maybe I'm overly cautious and the GPL is more lenient than this in reality.
We aren't talking verbatim generation of entire packages of code here, are we? Code snippets are surely covered under fair use?
...for "purposes such as commentary, criticism, news reporting, and scholarly reports"? Sure.
For a commercial product? Best check with your lawyer...
In general it is not fair use if you are using the material for the same scope as the original author[0] or if you are doing it just to namedrop/quote/omage the original.
It is possible to argue that a snippet can be too small to be protected, but that would not be because of fair use.
[0] Suppose that some Author B did as above and copied a snippet of code in their docstring to exlain buggy behaviour of a library they were reimplementing. If you are then trying to reimplement B's libary you can copy the same snippet B copied, but you likely cannot copy the paragraph written by B where they explain the how and the why of the bug.
I guess the likelihood decreases as the code length increases but the likelihood also increases the more constraints on parameters such as code style, code uniformity etc you pose.
If someone somehow managed to do that and then happened to have accidentally copied someone's code, how believable would their argument be?
No, and humans who have read copyrighted code are often prevented from working on clean room implementations of similar projects for this exact reason, so that those humans don't accidentally include something they learned from existing code.
Developers that worked on Windows internals are barred from working on WINE or ReactOS for this exact reason.
That's just copying with extra steps.
The way to do it legally is to have 1 person read the code, and then write up a document that describes functionally what the code does. Then, a second person implements software just from the notes.
That's the method Compaq used to re-implement the original PC BIOS from IBM.
wasn't it to have one person run tests of what happened when different things were done, and then write up a document describing the functionality?
In other words I think one person reading the code is still in violation?
>by reverse engineering and then recreating it without infringing any of the copyrights associated with the original design.
reverse engineering is not 'reading the code'.
Except that with AI we can more easily (in principle) provide provable provenance of training set and (again in principle) reproduce the model and prove whether it could create the copyrighted work also without having had access to the work in its training set
Maybe, but that would still be copyright infringement. See My Sweet Lord.
Not necessarily. If it's just a small snippet of code, even an entire function taken verbatim, it may not be sufficiently creative for copyright to apply to it.
Copyright is a lot less black and white than most here seem to believe.
Copilot similarly isn’t the one checking in the code. So it’s on each user. That said, Copilot at some point probably needs to add some type of copyright detection heuristics. It already has a suppression feature, but it probably also needs to have some type of checker once code is committed and at that point Copilot generated code needs to be cross-referenced against code Copilot was trained on.
But only snippets as far as I can tell.
This is the codeexample linked from the author:
https://web.archive.org/web/20221017081115/https://nitter.ne...
It is still not trivial code, but are there really lot's of different ways on how to transpose matrixes?
(Also the input was "sparse matrix transpose, cs_", so his naming convention especially included. So it is questionable if a user would get his code in this shape with a normal prompt)
And just slightly changing the code seems trivial, at what point will it be acceptable?
I just don't think spending much energy there is really beneficial for anyone.
I rather see the potential benefits of AI for open source. I haven't used Copilot, but ChatGPT4 is really helpful generating small chunks of code for me, enabling me to aim higher in my goals. So what's the big harm, if also some proprietary black box gets improved, when also all the open source devs can produce with greater efficency?
If a medical procedure is proven to be life-saving, what happens worldwide? Doctors are forced to update their procedures and knowledge base to include the new information, and can get sued for doing something less efficient or more dangerous, by comparison.
If you write the most efficient code, and then simply slap a license on it, does that mean, the most efficient code is now unusable by those who do not wish to submit to your licensing requirements?
I hear an awful lot of people complain all the time about climate change and how bad computers are for the environment, there are even sections on AI model cards devoted to proving how much greenhouse gases have been pushed into the environment, yet none of those virtue signalling idiots are anywhere to be seen when you ask them why they aren't attacking the bureaucracy of copyright and law in the world of computer science.
An arbitrary example that is tangentially related: One could argue that the company sitting on the largest database of self-driving data for public roads is also the one that must be held responsible if other companies require access to such data for safety reasons (aka, human lives would be endangered as a consequence of not having access to all relevant data). See how this same argument can easily be made for any license sitting on top of performance critical code?
So where are these people advocating for climate activism and whatever, when this issue of copyright comes up? Certainly if OpenAI was forced to open source their models, substantial computing resources would not have been wasted training competing open source products, thus killing the planet some more.
So, please forgive me if I find the entire field to be redundant and largely harmful for human life all over.
The problem here is that Microsoft is effectively saying, "copyright for me but not for thee." As long as Microsoft gets a state-enforced monopoly on their code, I should get one too.
If you don't "slap a license on it" it is unusable by default due to copyright.
This. People seem to forget that generative AIs don't just spit out copyrighted work at random, of their own accord. You have to prompt them. And if you prompt them in such a way as to strongly hint at a specific copyrighted work you have in mind, shouldn't some of the blame really go to you? After all, it's you who supplied the missing, highly specific input, that made the AI reproduce a work from the training set.
I maintain that, if we want to make comparisons between transformer models (particularly LLMs) and humans, then the AI isn't like an adult human - it's best thought of as having a mentality of a four year old kid. That is, highly trusting, very naive. It will do its best to fulfill what you ask for, because why wouldn't it? At the point of asking, you and your query are its whole world, and it wasn't trained to distrust the user.
If you, not I, uploaded my GPL'ed code to Github is the blame on you then?
Definitely not me - if your code is GPL'ed, then I'm legally free to upload it to Github, and to an extent even ethically - I am exercising one of my software freedoms.
(Note that even TFA recognizes this and admits it's making an ethical plea, not a legal one.)
Github using that code to train Copilot is potentially questionable. Github distributing Copilot (or access to it) is a contested issue. Copilot spitting out significant parts of GPL-ed code without attaching the license, or otherwise meeting the license conditions, is a potential problem. You incorporating that code into software you distribute is a clear-cut GPL violation.
If we think of Copilot as a (de)compression algorithm plus the compressed blob that the algorithm uses as its database, the algorithm is fine but the contents of the database pretty clearly violate GPL.
If I start creating a car by using a blueprint of Fords to create something at what point will it be acceptable? I'd say even if you rework everything completely Ford would still have a case to sue you. I can't see how this is any different. My code is my code and no matter how much you change it, it is still under the same licence as it started out with. If you want it not to be then don't start with a part of my code as a base. In my opinion the case is pretty clear: This is only going on because Microsoft has lots of money and lawyers. A small company doing this would be crushed.
In a lot of scenarios, there is an existing best practice or simply only one real 'good' way to achieve something - in those cases are we really going to say that despite the fact a human would reasonably come to the same output code, that the AI can't produce it because someone else wrote it already?
It's not a problem in practice. It only does so if you bait it really hard and push it into a corner, at which point you may just as well copy the code directly from the repo. It simply doesn't happen unless you know exactly what code you're trying to reproduce and that's not how you use Copilot.
It should be easy to check Copilot's output to make sure it's not copied verbatim from a GPL project. Colleges already have software that does this to detect plagiarism.
If it's a match, just ask GPT to refactor it. That's what humans do when they want to copy stuff they aren't allowed to copy, they paraphrase it or change the style while keeping the content.
https://github.blog/2022-11-01-preview-referencing-public-co...
If the data could only used by other open source projects, e.g. open source AI models, I don't think anyone would complain.
You could argue "well, but anyone can use the code on Github" and while that's technically true, it's obvious that with both Github and OpenAI being owned by Microsoft, OpenAI gets a huge competitive advantage due to internal partnerships.
People put code on github to be read by anyone (assuming a public repository), but the terms of use are governed by the license. Now you've got a system that ignores the license and scrapes your data for its own purpose. You can pretend it's human but the capabilities aren't the same. (Humans generally don't spend a month being trained on all github code and remember large chunks of it for regurgitation at superhuman speeds, nor can they be horizontally scaled after learning.)
You can still be of the opinion that this is fine, and I may or may not be fine with it as well, I just don't think the stated reason holds up to logic and other opinions ought to "baffle" you
> Copilot does not 'steal' or and reproduce our code - it simply LEARNS from it as a human coder would learn from it.
Not "the terms of use you agreed to allow them to do it". Different argument with different amount of merit in my opinion
For MIT licenses that's impossible currently because of the requirement to mention the authors.
Personally I have eschewed any personal use of Github since the MS aquisition and only ever use it where that's mandated by a client (so not my code). If you clone my code from elsewhere into a Github repo, that's just rude and contrary to me every intent and wish.
I think it's time to add a "No GitHub" clause as an optional add-on to the various open-source licenses.
Intellectual property is a nebulous concept to begin with, if you really try to understand it. There's a reason copyright claim systems like those at YouTube don't really concern themselves with ownership (that's what DMCA claims are for) but instead with the arbitrary terms of service that don't require you to have a degree in order to determine the boundaries of "fair use" (even if it mimics legal language to dictate these terms and their exemptions).
The problem isn't AI. The problem is property. Ever since Enclosure we've been trying to dull the edges of property rights to make sure people can actually still survive despite them. At some point you have to ask yourself if maybe the problem isn't how sharp the blade you're cutting yourself is but whether you should instead stop cutting. We can have "free culture" but then we can't have private property.
Also, AI doesn't learn "like a human". Neural networks are an extremely simplistic representation of a biological brain and the details of how learning and human memory works aren't even all that clear yet.
Open source code usually comes with expectations for the people who use it. That expectation can be as simple as requiring a reference back to the authors, adding a license file to clarify what the source was based on, or in more extreme cases putting licensing requirements on the final product.
Unless Microsoft complies with the various project licenses, I don't see why this is antithetical to the idea of open source at all.
If I took two copy-righted pictures and layered them on top of each other at 50% opacity. Would that be OK or copy right infringement?
AI models just use more weights/biases and more images (or any input).
if(nonfree_software){
// unhappy path
}
If I scanned a thousand polaroid pictures, and took their average RGB values and created a LUT that I could apply to any photograph to make it look "polaroidy" - would that be learned? Or the application of a statistical inference model? This alone is probably far enough abstracted to never be an ethical or legal issue. However, if I had a model that was only "trained" on Stephen King books, and used it to write a novel, would that be OK? Or do you think it would be in the realm of copyright infringement?
By your definition anything a computer does means it has learned it. If I copy and paste a picture, has the computer "learned" it while it reads out the data byte-by-byte? That sure sounds like it is "studying" the picture.
"AI" and "ML" are just statistics powered by computers that can do billions of calculations per second. It is not special, it is not "learning". To portray some value to it as something else is disingenuous at best, and fraud at worst.
In your stephen king example I would say it's still learned, because the "code" is a general language model that can learn anything. It's just you decided to only train it on stephen king novels. If you have an image model that trained 100% on public domain images and finetune it to replicate a specific artist's style I would personally think the finetuned model and its creator is maybe violating copyright.
But when it comes to learning I would say when you write a program whose purpose is to learn the next word or pixel, but it's up to the computer to figure out how to do that, the computer is learning when you feed it input data. It's the program's job to figure out the best way to predict, not the programmer. (it's not that black and white given that the programmer will also sometimes guide the program, but you get the idea)
When you write a program that does one or several things, it's not learning.
I think it's something to do with the difference between emergent behavior from simple rules and intentional behavior from complex rules.
If I created a program to read words from the input and assign weights based on previous words, I could feed in any data. Just like the polaroid example. (I suggested that the polaroid example was abstract enough not to be an ethical/legal problem because I believe it is mostly transformative, unless the colours themselves were copyrighted or a distinct enough work in themselves.)
Now If I only feed in Stephen King books and let it run, suddenly it outputs phrases, wording, place names, character names, adjectives all from Stephen King's repertoire. Is this a 'general language model'? Should this by copyright exempt? I don't think this is transformative enough at all. I've just mangled copyrighted works together, probably not enough to stand-up against a copyright claim.
I think people use AI and ML as buzzwords to try and obfuscate what's actually happening. If we were talking about AI and ML that doesn't need training on any licensed or copyrighted work (including 'public domain') then we can have a different conversation, but at the moment it's obscured copyright theft.
If you train a model to replicate 10000 specific artists, I could also get behind it being more like theft.
But if the intention was to train with random data (and some of it could be copyrighted) just like your polaroid example to generate anything you want, I'm not so sure anymore.
I feel the intent is the most important part here. But then again I don't know the intent behind these companies, and I guess you don't either. Maybe no single person working in these companies know the intent either.
It also gets murky when you have prompts that can refer to specific artists and when people who use the models explicitly try to copy an artists style. In the case of stable diffusion, if the CEO's to be believed the clip model had learned to associate images of greg ruktowski and other artists to images that were not theirs but in a similar style[0]
Even murkier is when you have a base model trained on public data, but people finetune at home to replicate some specific artist's style.
[0] https://twitter.com/EMostaque/status/1571634871084236801
You wouldn't. LUT would.
Since it’s data that should be cool right ?
In my mind (and I suspect others too) in machine learning context, statistical inference and learning became synonymous with all the recent development.
The way I see it, there's now a discussion around copyright because people have different fundemental views on what learning is and what it means to be a human that don't really surface.
copyright is a thing, AI do not change that.
> does not 'steal' or and reproduce our code - it simply LEARNS from it as a human
And here we have the central problem, does it act like a human or does it not act like a human? Humans copy things they learn all the time, some of us know various songs by heart, others will even quote entire movies from memory. If AI can learn and reproduce things like humans do then you need to take steps to ensure that the output is properly stripped from any content that might infringe on existing copyrighted works.
In general individuals will have problems with the first L of LLMs - unless the community invents a way to democratise LLMs and deep learning in general. So far deep learning space a much less friendly place for individuals than software was when ideals of open source movement were formed.
There are multiple open source LLMs out there that can be extended.
We can already see it in AI art scene. People are training their own checkpoints and LoRAs of celebrities, art styles and other stuff that aren't included in base models.
Some artists demand to be excluded from base model training datasets, but there's nothing they can do against individuals who want to copy their style - other than not posting their art publicly at all.
I see the same thing here. If your source code is public - someone will find a way to train an AI on it.
I thought that was a sarcastic remark, given the capitalization of 'learn', but followed by IMHO dispelled that part.
We have no idea how humans learn, and the 'AI' has a statistical approach, not much more than that.
Have you personally ever put out something in public domain?
I don't really want this to comment to be perceived as flame bait (AI seems to be a very sensitive topic in the same sense as crypto currency), so instead let me just pose a simple question. If Copilot really learns as a human, then why don't we just train it on a CS curriculum instead of millions of examples of code written by humans?
The problem is that in reality, even though the original data is gone, a language model like Copilot _can_ reproduce some code byte for byte somehow drawing the information from the weights in its network and the result is a reproduction of copyrighted work.
To say "it's not a database, it's a language model, and that means it extracts generalized patterns from viewing examples, just like humans" to me that just means that occasionally humans behave like language models. That doesn't mean though that therefore it thinks like a human, but rather sometimes humans think like a language model (a fundamental algorithm), which is circular. It hardly makes sense to justify that a language model learns like a human, just because people also occasionally copy patterns and search/replace values and variable names.
To really make the comparison honest, we have to be more clear about the hypothetical humans in question. For a human who has truly learned from looking at many examples, we could have a conversation with them and they would demonstrate a deeper sense of understanding behind the meaning of what they copied. This is something a LLM could not do. On the other hand, if a person really had no idea, like someone who copied answers from someone else in a test, we'd just say well you don't really understand this and you're just x degrees away from having copied their answers verbatim. I believe LLMs are emulating this behavior and not the former.
I mean, how many times in your life have you talked to a human being who clearly had no idea what they were doing because they copied something and didn't understand it all? If that's the analogy that's being made then I'd say it's a bad one, because it is actually choosing the one time where humans don't understand what they've done as a false equivalence to language models thinking like a human.
Basically, sometimes humans meaninglessly parrot things too.
This just means it's a really efficient lossy compression algorithm, not that it learns like a human.
I've never studied computer science formally but I doubt students learn only from the CS curriculum? I don't even know how much knowledge CS curriculum entails but I don't for example see anything wrong including example code written by humans.
Surely students will collectively also learn from millions of code examples online alongside the study. I'm sure teachers also do the same.
A language model can also only learn from text, so what about all the implicit knowledge and verbal communication?
A CS graduate could workout how to write software without doing that.
So they’re just pointing out the difference in “learning”.
But I'm saying there's a big difference between a CS graduate and some current LLM that learns from "the CS curriculum". A CS graduate can ask questions, use google to learn about things outside of school, work on hobby projects, study existing code outside of what's shown in university, get compiler feedback when things go wrong, etc.
All a language model can do is read text and try to predict what comes next.
Does it though? It "learns" correlations between tokens/sequences. A human coder would look at a piece of code and learn an algorithm. The AI "learns" token structure. A human reproducing original code verbatim would be incidental. AI (language model, at least) producing algorithm-implementing code would be incidental.
This is really not how LLMs work.
What humans do to learn is intuitive, but it is not simple. What the machine does is also not simple, it involves some tricky math.
Precisely if the process was simple, then it could be more easily argued that the machine is "just copying" - that is simple.
There's a lot of nuance here.
What the machine is doing "looks similar to what humans do from the exterior", the same way that a plane flying "looks similar" to a flying bird. But the airplane does not flap its wings.
> kind of irrational and antithetical to open source ideas
Open source ideas are not the only ideas in town.
It doesn't LEARN anything, let alone like a human coder would. It has absolutely zero understanding. It's not actually intelligent. It's a highly tuned mathematical model that predicts what the next word should be.
Rather, Copilot is a tool. Microsoft/ClosedAI operate this tool. Commercially. They crawl original works and through running ML on it automatically generate and sell derivative works from those original works, without any consent or compensation. They are the ones who violate copyright, not Copilot.
> it simply LEARNS from it as a human coder would learn from it.
The LLM doesn’t learn, the authors of the LLM are encoding copyright protected content into a model using gradient decent and other techniques.
Now as far as I understand the law, that’s OK. The problems arise when distribution of the model comes into play.
I’m curious, are you a programmer yourself? Don’t take this the wrong way, but I want to understand the background of people who coming to the kind of conclusion you seemed to arrive at about how LLMs work.
As an aside, it seems really strange to invoke "open source ideas" as an argument in favor of a for-profit company building a closed source product that relies on millions of lines of open source code.
Here is Microsoft doing as Microsoft does…
You may be right that this is antithetical to "open source" ideas, as Tim O'Reilly would've defined it - a la MIT/BSD/&c., but it's very much in line with copyleft ideas as RMS would've defined it - a la GPL/EUPL/&c. - which is what's being explicitly discussed in this article.
The two are not the same: "open source" is about widespread "open" use of source code, copyleft is much more discerning and aims to carefully temper reuse in such a way that prioritises end user liberty.
Both claims here are incorrect even though pretty much everyone gets this wrong. When someone who has obtained the right to redistribute some code only under the GPL uploads it to GitHub, that person (and that person only) violates the terms of the GPL. The GPL requires further redistribution only under its own terms, but uploads to GitHub come with a grant of a too-permissive license that a GPL licensee does not have the right to grant.
When GitHub proceeds to use the uploaded code to train copilot, they (probably) are abiding by the terms of this new license they have been (fraudulently) granted. They are not bound by the GPL, that's not how licenses work: they've got the other one. Now, GitHub has a big weakness here which is that they ought to know they're being granted licenses that the putative licensors have no right to grant. But that still would not make them in violation of the GPL, just of the original copyright.
but they are bound by the original copyright and are in violation of it
Do you actually have the necessary rights to upload some else's code to github?
When you upload to github, you give it special rights not merely for redistribution and CI stuff.
You give Github the rights to use the code for other github projects. That alone might not be compatible with some licenses (think GPL virality).
So if your software is BSD or anything without attribution I can probably upload it without problems.
But if your license requires so much as attribution, can I give Github the rights to use the code for any other internal project they might have?
Remember that in Github TOS you give grants for any github service, and in some special cases it requires this to be without attribution.
IANAL but I think I lack the necessary rights here, regardless of copilot.
BSD and MIT licenses absolutely do require an attribution.
To your point about redistribution, see this clause in D.4: “This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service, except that as part of the right to archive Your Content, GitHub may permit our partners to store and archive Your Content in public repositories in connection with the GitHub Arctic Code Vault and GitHub Archive Program.”
But section D.7 ("Moral Rights") is about waiving moral rights, which at least in Italy is not legal, and then says: "To the extent this agreement is not enforceable by applicable law, you grant GitHub the rights we need to use Your Content without attribution".
It mostly seems like poor wording of that section not being limited to archiving/displaying, but there is a case where you grant Github something without need of attribution.
Which is against most Open source licenses, which means that I can't upload someone else's code.
Again IANAL, and theoretically open source projects want maximum distribution, but I think this does mean that technically I would grant to Github something I can't grant.
Then what gives them the right to use the code for creating proprietary software (Copilot)?
First, LLMs learn patterns, not just copy and paste. If they generate verbatim copies of any non-trivial-enough part that would be the subject of a copyright license, yes, it would be copyright infringement. Yet, could anyone give such practical examples? And if so, how do they differ from a software engineer who copies and pastes code?
Second, if the code is hosted anywhere else, there is no guarantee that Copilot (or another model) won't learn from that. The only way to make sure no one and nothing will learn from open-source code is to make it as closed as possible.
Third, for me, the crucial part of open-source code is maintenance. GitHub is there and works well both as a platform for creation (I consider GitHub the most productive social network) and an archive. "No GitHub" (even as a mirror) means that the code is likely to be stored in places less likely to engage collaborators and less likely to last long.
Yes, there are many such examples.
https://twitter.com/docsparse/status/1581461734665367554
https://twitter.com/mitsuhiko/status/1410886329924194309
https://codeium.com/blog/copilot-trains-on-gpl-codeium-does-...
It is mutatis mutandis the same but is that a problem? I'm sure many would say so, I'm not convinced.
Ultimately if his code is out there a Google search could bring up a snippet without the license visible and I might copy paste that. The crux is the same code might be presented without context.
Copilot is just a tool and the personal responsible for it's safe usage is the human behind it.
In my world view, if I copy a picture off Google image search ultimately I am morally the one who infringed copyright not Google.
I have an idea why, but... why exactly? What about a web scraper (that I made, similarly to that of Google) that downloads images? What if it is randomly downloading images and not intentionally a specific one?
https://twitter.com/docsparse/status/1581461734665367554
It implements common operation in a standard way as far as I can tell. AI re-implementation is not identical.
It seems more like multiple discovery than plagiarism let alone re-producing a copy without a license.
It's always the same two examples, and I would not classify that as "many", especially since that fast inverse square root function has been shown to be on GitHub and other sites countless of times with all sorts of different licences (which is wrong, but copilot doesn't seem to do better or worse than humans in this regard).
That codeium.com is just asking leading questions, or the AI equivalent of that.
They differ because author of the code did not agreed to use his code to teach LLM.
> Second, if the code is hosted anywhere else, there is no guarantee that Copilot (or another model) won't learn from that
This issue should be resolved not by author of the code, but by Copilot or any other LLM team.
> Third, for me, the crucial part of open-source code is maintenance.
Even if author is against it? This is a valid argument, but results is not very different from pirating of software.
When to the third point - well, it is up to the author, and I respect that (regardless if I would do the same thing). People have the right to not share it at all, or share it as a copyrighted piece of software, or with any other limitations. Though, all limitations (and copyleft is a limitation) affect its usage.
For example, GPL controls what kind of projects can use my source code. Maybe there could be an addendum to GPL that requires all LLMs trained on the source code to be open source. Sure, that won't guarantee that Copilot-like bots won't be trained using my code. But it does give me a legal framework to stop big corporations from profiting off such Copilot-like bots without making them open-source as well.
> Is this a legal document?
> No, it isn’t. If the project is under an open source license, it means that everyone can share a copy – even on GitHub – of the licensed material under certain conditions. A license restricting this right wouldn’t be open source anymore. However, since GitHub may not respect the terms of licensed code that is hosted on their servers, not uploading the code of others there is, in fact, an ethical choice.
emphasis mine. It's a "please be nice", not a "I want to enforce things"
I respect developers right to put any restrictions on the code they share with the world. But I believe it should either be explicit or not restricted at all. Either write the license that says exactly what you want or otherwise don’t shame people into the desired behavior.
Edit: one could even add a more generic statement to the license stating that it’s forbidden to share the code on any platform that would use it to train their AIs per their ToS, so you don’t need to single out GitHub and potentially others in the future.
it is a wish. if someone says "please don't wear shoes in my home" i hope you would honor their very simple and understandable personal wish without setting up a contract for it?
i mean, just be a bit more human, please.
Stallman would prefer I not use any closed source software to read his blog, including OS, drivers, web browser, etc.
People routinely ignore unreasonable requests. Asking me to not wear shoes in your home is reasonable. Asking me not to give a copy to Joe after telling me I can give a copy to whomever I want is unreasonable.
We already established no one is forcing you and if you don't respect the author you don't get to be respected for your decision to ignore them and will earn snarky remarks. (and rightfully so, in my opinion)
This is a repetition of your claim, not an argument for it. Counterpoint: It's entirely reasonable to ask you not to give a copy to Joe.
(1) Though I don't know if this has been actually tested in court - courts in India have more freedom to broadly interpret social contracts like the GPL, unlike the US courts, and a positive outcome in favour of upholding the license even in such cases could be possible).
I disagree here. The idea, intent, philosophy is one (crucial) thing, the resulting practical artefact (here the license) is another. It works exactly as it was designed to work.
People/companies modifying GPL software for their own use (internal or external) without redistributing the software itself (so without requirement to redistribute the code) existed before SAAS grew on, only at the time, the small scale of this made it a bargain that was "interesting" only depending on one's capacity/hubris to maintain an internal fork on their own.
*aaS hugely tipped the scale, and the side effects, but the mechanics are the same.
And yes, that may not have been the original intent, and the AGPL is as valid a license as a reaction to provide a new tool more in line with the original intent, but that doesn't make the use of the existing GPL all within what it actually enables anyone to, invalid or unethical.
(but maybe only in a specific perspective of the framework of the original intent)
How is this case different?
It took me some time to get that this is a generic call (to be followed and reused), still with no clear ownership, rather than a specific claim to a specific code/project (or is it?).
>> No, it isn’t.
My limited understanding of international copyright law is that copyright disclaimers/licenses per Berne Convention are documents stating the author's wishes no different than this one. When a judge is going to be shown this page where the author states his explicit wish to not have his code uploaded code on Github, in direct conflict with his wishes stated in the LICENSE.txt file, I don't see how it will hold up in court as free software code.
Also don't forget to go back in time to let your former self know which hosting providers will violate your license.
I mean even doing that will not protect anybody's code, Microsoft doesn't care, they might also be scanning gitlab or bitbucket and doing model training on these. The only way to protect your code at that point is to stay closed source. It's as simple as that.
It's no different from all these models trained on content without permission, whether it be articles or photos,... until all these corporations are sued or copyright laws change to adapt ML, they'll train their model on content regardless of the its license since they are getting away with it.
It's pretty clear in the article, I hesitate just just cut and paste it here, but the idea is that Github is not respecting some terms of some licenses, so if you upload another authors code to github, you are exposing the author rights to github's depradations.
I agree with 'can', but I find 'should' a weird choice of words here.
I don't think we are better off by phrasing hyper-specific licenses (or laws).
Github has been consistently arguing that they're not bound by the distribution license when training their Copilot model, so I'm not sure what difference that would make.
This is even true for e.g. MIT licenses as while they are very permissive they still require a form of attributions which Copilot doesn't provide.
Also for anyone arguing that such ML models "lern", at least the moment they are not exatcly perfect sized or sized below that they are basically guaranteed to verbatime encode partial copies of code, it's fundamental consequence of their design. And while this encoding is "encoded" in some way it's not transformative in the definition of copyright law AFIK. I.e. any big model net is guaranteed to commit hard to trace copyright infringement.
> This is even true for e.g. MIT licenses as while they are very permissive they still require a form of attributions which Copilot doesn't provide.
those two things have nothing to do with each other. MIT gives me the license to upload. Copilot NOT giving me the license is irrelevant because the generated code isn't distributed under under an opensource license and thus has no relevance to the discussion of MIT or any other Open Source license.
additionally, requiring attribution and having the right to upload it are separate things. A license MAY require attribution and it may not. Rights can be granted without such a thing. See Public Domain, and CC0 for examples.
> generated code isn't distributed under under an opensource license
you still have to comply with licenses and if code pilot spits out code which contains nearly verbatim code which was MIT licensed it is not copilot which needs to grant you the license but the original code owner (or you need to at least properly attribute it, through in case of e.g. GPL that is not enough)
but through the act of emitting code via code pilot github does distribute code (without proper attribution) and they do so by getting rights from the uploader to be able to do so. Except that most people uploadig do not have such rights for anything which isn't their code.
People must separate their fascination with tech as such, from the predominant tech business models that are basically - you can't put lipstick on a pig - parasitic.
you could proactively scan github for your code and try to get them to purge it if you find it i suppose, if that would remove your code from copilot. but even that is not a great solution because you would need to prove you're the actual author, and github would probably need to be involved in building a mechanism to do so, but they don't give a shit.
i think the reality is the LLMs have eroded copyright protections and trying to fight it isn't likely to pay off
https://www.bigcode-project.org/docs/about/the-stack/
I am hopeful that we can make the removal request form the standard industry practice.
This generally is the whole point of open sourcing code: to allow others to republish with or without modifications. If you don't want that, don't open source your code. Github is a perfectly valid and legal place to publish code. Lots of people have done so.
The jury is still out on whether using open source code as training data is fair use or constitutes a copyright violation. But the safest assumption at this point is that it is a combination of fair use and hard to prove that any violation happened to begin with (i.e. good luck stamping your feet in anger in a court room).
Which will no doubt make for some fun court cases but will also take many many years to come to any kind of conclusion. I would recommend not to get your hopes up on that. Historically, copyright law has not been updated a lot to deal with any kind of technical change. And that's just inside the US. The world is bigger than that of course. A safe assumption is that judges will seek to interpret any new cases under centuries of existing interpretations. Because that's what they do. Any changes to the law would have to go through politicians. And then the interpretation of that relative to other laws is again up to judges. Which is why we have such wonderful things as DMCA and a few other attempts to close loopholes in copright law.
Of course time is not your friend here. By the time this rolls through the courts some years down the line, existing practice will have evolved to be hopelessly dependent on language models for just about anything. Including interpreting the law (e.g. chat gpt seems to be acing exams) and doing all sorts of complicated things in science, engineering and technology. So, any legal outcome that that is somehow illegal is going to be combination of unpopular, economically damaging, and therefore unlikely. Think big corporations ki
Ok, not a lot, but it has been updated. Chip mask is a good example: https://en.wikipedia.org/wiki/Semiconductor_Chip_Protection_.... If language model gets as important as chip mask, I don't see why similar update wouldn't be made.
Most of the people complaining seem to be a narrow subset of open source developers that like open source but not people doing anything for profit with their source code. I.e. not your typical multi billion dollar companies that are worried about their IP getting stolen. So, not a lot of lobbying power or ability to get organized.
If it is that simple then clearly Copilot need to include the licence of the code when they republish it. Changing GPL'ed code doesn't make it not-GPL code at any point. If it started out as GPL it is still GPL later in the line when a user sees it in Copilot.
It's one thing when individual devs or small teams don't respect a software license, ideally everyone would understand and follow them but that's not realistic. When Microsoft damn well understands the licenses and ignores them to train an AI and sell a product, that's pretty damn egregious.
Given the current state of the law and how licenses are written it's impossible declare that "yes, this is absolutely illegal". People may not like it, and (arguably) it may go against the spirit, but that is also a very different thing. Right now the situation is just ambiguous and unclear.
All of this absolutely matters, because the solution to "Microsoft is acting illegally" is a lawsuit (which, I believe, is already underway) while the solution to "the license and law is unclear" is different licences and/or laws. The solutions are completely different.
This is a reason for Microsoft (which you have given to them) that explains why they're ignoring the license, and why you think it's fine. It's a common refrain: "Everything is so complicated." It doesn't change the fact that they're ignoring the license.
When Copilot inserts a code snippet that matches perfectly with one of the samples it was trained on, and that sample had a copyleft license, how is that not a license violation? Or is the main argument that copyleft licenses aren't enforceable in general?
If I own "the thing", I want you to pay for each new, distinct kind of use for it. But if it's your thing I want to use, I want you to have a permissive license.
After you have cleaned everything off of GitHub and separated yourself from the ecosystem it (AI bots) will be everywhere and your code gobbled up again.
The only thing that is worth a fuck is a working product. Not your boilerplate or fancy snippets that you want to claim ownership over so as to stop the evil Microsoft from benefitting the community.
Are we devs or luddites?
Most open-source projects are not valuable, and even if they are popular, they could be easily replaced in a few days by another programmer. And usually the license doesn't matter, because nobody forks the project anyway, all development is done in a central repository and the project owner signs off on all of it
Even without this, in terms of copyright, since Copilot doesn’t do what your public code does, and it only uses your code to train, it is a transformative use, and would be fair use. It’s possible that a court case will find otherwise, but I think that’s unlikely. The only case I think it will become disallowed, is if Congress passes a law about it.
If Congress does pass such a law, GitHub’s market power in this domain only goes up, since the EULA gives it the covenants.
Unless you can reproduce a substantial portion of a repo, I think it’s going to be an uphill battle to argue it isn’t fair use. Though I suspect Copilot’s suppression feature will make doing so impossible.
Either way, the whole copilot thing smells of 'the issue of copyright infringement is more copyright infringement' to me.
(And I don't see why we cannot add a clause saying that the source code is only meant for human developers and use of the source code in any machine learning system or to train any AI systems is prohibited without explicit case-by-case permission).
What many people actually want is for LLMs to respect Open Source licenses, and propagate those licenses to the derived works they create.
So? Maybe making so much code open source without any restrictions was a mistake in the first place. I know that I don’t want trillion dollars megacorps benefiting from my free open source code in any way. That would include LLM training.
The blocks retrieved are very small and many of them occur frequently. “if” followed by “(“ for example — hardly worthy of copyright, but we also know that they were literally taken from copyrighted material.
(I don’t think the model starts out with any existing knowledge of syntax / grammar of, say, Python?)
Even if some of that material was public domain, a lot of it wasn’t and at best requires attribution; at worst, full licensing conditions.
To put it another way: it doesn’t matter how many vegetable ingredients they throw into the sausage or how elaborate the sausage making machine is: if they put pork in, the stuff that comes out ain’t vegetarian.
I'm not sure about this. In most cases the result would be similar to a human reading a lot of opensource and later when writing use patterns that they'd learned. It's only in the edge cases where there's clear 'plagiarism' on a niche prompt that it would be problematic. An more direct solution isn't to take everything off GitHub, but rather to not allow Copilot to do near-literal copy/paste.
If we moved opensource to BitBucket, there's no protection that it wouldn't do the same as Copilot. Attack the problem directly.
A way to think of this banner is that of signing a publicly visible petition to make Copilot behave as humans abiding to licenses do.
While encouraging people to not distribute code via Github may mitigate the issue some, the actual issue is how Github has mass-automated the process of violating open source licenses. Github should pay a fine for every suggestion Copilot produces that violates a software license, plain and simple. Don't blame the people that unknowingly upload code to the training dataset.
Weasel wording to evade legislation. It's not an unlicensed taxi, it's ride sharing. It's not an illegal hotel, it's couch surfing. It's not code licence infringement because it's learning.
It's lawyers finding loopholes for finance to avoid expenses that gives them an edge over the poor suckers who play by the rules. The tech is just a tool to this purpose.
And most people forget TANSTAAFL. The costs for the cars and their infrastructure, the load tourists put on a place, and the effort for writing the code are still there. The "innovators" just found a way to make somebody else pay for what should be their cost center.
The tone on one of these tools is hypocritical. When it comes to digital art general sentiment is that it's inevitable and artists need to up their game. This sentiment is not being repeated for code generators.
There's a miss match between producers of the content used to create the models and those using it. At the moment its mostly software developers who are using the output of Copilot, whilst its mostly non artists who are using MidJourney/StableDiffusion.
Learning from source code is like learning from Photoshop files or Logic projects. Artists usually do not share them. In music, not even the labels get anything than the final mixdown (the music equivalent of the binary); the stems or projects remain with the studio - just like the source code in software projects tends to remain with the studios.
In software, we started showing and sharing source code under the assumption that it makes other humans making software better. People still rarely do this in music, as holding up industry secrets is deemed more important than fostering a community of personal growth.
We don't open source our code so an AI can grow. We do it so that humans can grow. There's no need to publish our code to distribute our final binary. This breaks the common understanding of why it's a good idea to share your source code, and as a result, people might become more protective of their code again.
In many ways, your comment is similar to all the comments saying "but you open-sourced it, surely you must know that people will then put it in their commercial code?"
Also the code is the end product. Similarly, artists don't share their project file because it’s not the point. This is like saying programmers not sharing their vim macros.
Apparently it's directly available in Logic (or was)
https://old.reddit.com/r/Logic_Studio/comments/gic2r2/logic_...
Hold on just a second here. How do you know what "people" "want"? Who put you in charge of speaking for them?
Yes, there have been some loud complaints, but given that millions of people are involved, the overwhelming majority of whom haven't expressed an opinion one way or the other, I think it's a bit premature for this kind of blanket statement.
Secondly, there's a difference between what people want and what people are legally entitled to. If you ask people if they want a million dollars, most of them will say "Sure!". That doesn't mean they're going to get it.
Existing copyright controls the making of copies. That's it. There's a fudge factor in there called "fair use" that controls whether or not something constitutes an infringing copy.
Whether AI training data falls within that or not is going to have to be decided in the courts or by some type of government action. It's clearly not an exact copy...the actual pixels in the original work aren't anywhere in the database. But is what is in there close enough to be considered a "derivative work"? I don't know, and neither do you.
Again, it's way too soon to be making blanket statements on the issue.
While it may differ in other countries, in the United States the purpose of intellectual property laws, as expressed in the Constitution, is to "To promote the Progress of Science and useful Arts by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries".
Enriching the copyright holders (on the rare occasions that actually occurs) is a secondary consequence, not the prime purpose.
Does AI "subvert" "promoting the Progress of Science and useful Arts"? I don't think so. Quite the contrary... I think it advances the progress of science and the useful arts, if anything.
It's pretty well established that a description of a copyrighted work is not protected, even an extremely-detailed description (see, for example, the way that Phoenix assigned one team to write an extremely-detailed description of the IBM PC BIOS, then gave that description to a second clean-room team that hadn't seen any of the actual source code. The second team then produced a clone of the BIOS that could be sold without paying IBM anything).
The data stored in these models seems more like a "description" rather than a "copy" to me -- though, of course, there's no guessing what a court or legislature will decide.
It does subvert that purpose to the extent that it makes some people no longer willing to share their works. The entire purpose of copyright is to encourage the sharing of works.
Absolutely, it may. That was my point that this is exactly to be observed and decided.
1. To showcase their work
2. To get feedback
Also, I don’t think it’s fair that we have to wait for the majority response before we can form an opinion of whether or not something is good for them. Do we have to ask millions of people if they like it if their health benefits, social security or left thumb gets removed before we can say that they definitely won’t like it?
I think it is pretty safe to say that artists don’t like having their whole livelihood overnight. (Inb4 get better jerbs)
this is a general problem in open source. And with obfusication in an AI it is just unfair. If AI / codepilot would reference the source in a nice way, then I woundn't mind.
This is absolutely not the reason and it's surprising to read this opinion stated as a fact.
First of all, not revealing multi-tracks doesn't prevent personal growth. Human ear is pretty good at discerning individual elements of a mix. It's all in the final product. If you can't learn from it or can't hear something, either it's an unimportant part of the mix, your ear is not good enough yet to distill any meaningful information from the balances of the mix in question, or it's just not a good mix.
Secondly, multi-tracks are not the source code. Putting it in relevant terms, it's just the assets. What's important are processes and artistic decisions made while processing them and combining them in their final form.
Not revealing stems/multi-tracks is not a matter of gatekeeping, it's a safety measure to prevent unauthorized use of individual composition elements and creation of "alternative" mixes which break artistic integrity of the piece.
Music is art (at least a considerable portion of it), and and artist needs as much control over the final product as possible. Mixing is a part of it.
If musicians' priorities were for the world to share in their musical indulgence not just in listening but also in like-minded creation, they would release as much of their mixing tools/backgrounds/stems/etc as possible. Since the capitalist framework is a foregone conclusion, they need to keep some secrets in their sauce to produce scarcity to justify a price for their labor so they can eat... Because the farmers and butchers all keep gates of their own kind, too :)
They do. For a price. When you know how to ask.
The thing is, if you really want to « like-minded create », you do not need, not even want their stems (you can find most of them anyway online, if only for learning).
If your thing is not « consumption » but creation, your own voice/process/way matters more (and is more fun) than starting from the studio separate tracks.
As for the music _business_ side of things, attention (hence scarcity management) on and availability/capacity to provide are the two main levers you have if you want to live from it. As with any other business actually.
I disputed a specific mistaken proposition, that "holding up industry secrets is deemed more important than fostering a community of personal growth". It's a mistake on several levels: there's no dichotomy as it's presented in this sentence, and the reasons for not providing multi-tracks and project files are different.
> If musicians' priorities were for the world to share in their musical indulgence not just in listening but also in like-minded creation
So? They have other priorities so this is not what they do. Teaching is an industry that's only an adjacent and derivative activity to the cultural sphere of arts. It's of no concern for the artist, unless they choose to capitalize on their skills/knowledge/experience.
I think you’re being a little too altruistic. When I publish code, it’s not about helping people grow. It’s so that others can see the code they are using and that if there is a bug or mistake, they can correct it. Hopefully they share that fix back, but it’s not required.
Only if Copilot was issuing compiled binaries. In this case the end result is the code.
When was this decided? I don't remember any licences or foundational essays saying this.
Copilot can provide very specific implementations of problems with "in the style of $NAME" prompts. Consider a case where you put your life's work as GPL licensed code, and someone can reproduce + adapt your highly optimized matrix multiplication code with a simple prompt, without license. Even if it's not a "textbook license violation", you're lifting one's GPL licensed function and landing it to your codebase without its license. If your code base is not licensed under the same GPL version (or later if the repo allows), it's both unethical and license breach at the same time. Adaptation of the code doesn't matter.
Same is true for more strict, source-available licenses. They are open source, but not open to be reused. What will happen if you put a function derived from a codebase with these strict licenses with or without knowledge? You're again in a dangerous grey area from both license and legal perspective.
The issue we discuss is neither straightforward nor simple to navigate. I left GitHub because of this, and may tag my repositories with this badge.
Open source means nothing if copyleft is taken out of the picture, and licenses are simply ignored.
This.
Since Copilot arrived I thought that most open source developers would be fine with it if github simply even tried to acknowledge the original licences in any way.
I personally would be ok with even a very indirect aggregate group based thing that should be no burden at all for github. They make a big list with everyone's name in it and call it the copilot contributors, and provide some kind of page for it, then when copilot spits out code, it includes a link to that page, and/or a user includes that in the credits/authors for their project like any other credited source.
No excuses about how impractical it would be to cite 3 other authors for every line of output.
But they don't even do that tiny bit. They don't try and fall short, they don't even try. But they still take the goods. The goods are already free, and yet they still manage to steal them.
Challenge accepted?
With the speed at which things are advancing....
I open source my code so that programmers and users can grow. I don't care if that programmer or user is a meatbag or a machine (or both! or neither!).
What I do care about is that the license terms of said code are respected. The vast majority of my code may be under something permissive like MIT or (lately) ISC, but I do make giving credit where credit's due a condition of using the code I've written, for good reason.
That's where tools like Copilot make a misstep: by ignoring the conditions I've placed on the use of my intellectual property. Plagiarism is plagiarism, regardless if it's a human or AI or dolphin or Martian or whatever doing it.
That's also where tools like Copilot differ from e.g. StableDiffusion. AI-generated art doesn't (usually) involve copying and pasting snippets of existing artwork into a new work the way Copilot has been demonstrated to do on multiple occasions.
(My other "problem" is that I can guarantee Microsoft will assert double-standards when it comes to Copilot infringing on e.g. the GPL v. Copilot infringing on Microsoft's own EULAs - and I really really really want to see that happen via someone tricking Copilot into ingesting the Windows source code and vomiting that into Copilot users' IDEs verbatim)
FTR I personally think that training commercial ML algorithms on bulk data produced by people that haven't released it to the public domain or licensed it for such use is ethically dubious (regardless of the legality) whether the data consists of code, art, metrics or anything else.
HN is actually one of the most homogeneous public communities I've seen.
And judging by how few downmodded posts are around, self censoring.
But then the next day a different post on the same topic could have an entirely opposite hivemind, often no discernible reason.
I'm not so sure this is a huge effect. I often express "incorrect opinions", and get some amount of downvoting or ridicule, but not usually to a significant degree. When I've had my comments dogpiled, I can see that it happened because of how I stated my opinion more than because of the opinion itself.
Sometimes the same story will be flag killed one day, front page the next.
Look in a thread where China is mentioned and you'll see the real face of many HN'ers - if they aren't removed before you see it. Luckily the MOD'ing is pretty good.
Unfortunately since most people aren’t experts on most things, this turns into a bunch of fugazi. People who are financially or academically successful in business or engineering or marketing think they understand politics too.
This is what happens with topics like China. I don’t care about China much. So I end up defending it because the negativity towards it by the Global North is staggeringly inconsistent.
This doesn't seem obvious to me at all. It seems not only plausible that this could be the case, but unsurprising.
Almost nobody here reads every story in the feed. Any given story on HN is going to be read by a collection of people that differs from story to story. And most of those people won't make any comment whatsoever. It seems unsurprising to me that a story on, say, capital punishment could get a large number of people who are passionate about that topic, and a story on pedophilia could get a large number of people who are passionate about that topic, and the majority of commenters in both could hold opposite opinions without even one of them being hypocritical.
Is it, though? My impression is that HN opinions on possible copyright issues are generally mixed for both, and mostly dismissive about the job replacement stuff for both.
Source: over 20 years of working with their products and seeing the way they treat everyone else
I've seen this question posted on HN so many times in defence of Copilot but I have yet to encounter people who fit what you're describing: people in general raise a lot of issues with MidJourney/StableDiffusion/&c. too.
HN is particularly defensive when it comes to Copilot because Copilot deals with sourcecode which is central to the daily vocation of a large portion of HN's readership. HN focuses on Copilot because they're subject matter experts in what Copilot trains on.
Also - as mentioned in siblings - while criticisms of training material sourcing by various ML models is common, there are quite a few details that makes Copilot's sourcing that little bit worse.
> it's inevitable and artists need to up their game
If this is your answer it seems you didn't even bother to read the title (never mind the article). There are two issues with Generative AI to discuss: input and output. Your answer addresses output - i.e. fear that what is produced may compete with humans. This article, and the vast majority of HN criticism of Copilot exclusively addresses input: objection to copyrighted works being used to train models without author consent. These are two entirely separate discussions.
There are probably not as many artists here, as there are coders. So the general opinion will be against all tools which will harm the coders ground, but more open for the benefits AI will bring for the laymen, when they must not fear harm from them.
My effortless attempt: OMIJIB(Only My Information Job Is Brilliant)? O-My-Jib
I can’t handle hearing another person say “HN is a big place”. Sure. It’s also incredibly obvious how homogeneous it is and how HN essentially forces you to toe the line.
I’m a leftist in tech. HN is not the place where you can publicly behave as an assertive leftist otherwise would.
There’s no real difference it is just wagon circling once a thing threatens somebody personally.
The tech community has been one of the big proponents of “data wants to be free” and “the entire idea of intellectual property is bs” for a long time and that’s suddenly reversed.
To be even more extreme, a _lot_ of people think if anything is public, it is fair gain to consume and reuse. If you didn’t have to worry about getting paid, would you care about someone enjoying and using your work? Perhaps the opposite - that would be your motivation.
Such a license would not be considered "open" by the OSI because it imposes onerous conditions.
Even if they were made, I wonder how much such licenses would be used. If the permissive people don't like copyleft because of the conditions it imposes, they must hate this. Likewise the copyleft folks want freedom enforced, not taken away.
I don't see the point of this because quite frankly - if you want to prevent AIs using your data, you already lost that battle in the moment you uploaded your Code to the publicly accessible internet.
Did you write the absolute best implementation of X? I want to see it everywhere. Everywhere. In every single place where X is needed or discussed. Where all fine X are sold. Don't you? Or do you genuinely want to narrow the people who see your impl by some signifnicant percentage because you got frustrated with the ugly capitalism of a recent distribution mechanism?
If I invented Golden rice[0] I'd like to think I'd allow it to be sold at Panda Express and all the other evil capitalist chains, not just the local co-op whose business practices I prefer.
This may seem an extreme comparison, but to one who deeply cares about Free Software as defined by the FSF, proprietary software is bad. That's the difference between people who prefer Apache/BSD/MIT licenses versus those who prefer the (A)GPL.
Fighting fire with fire is one option. Fighting fire with water is another.
Yes, I know that it can occasionally break license by producing licensed code verbatim. But in my almost 2 years of using it daily I have never seen it happen first-hand, and I don't see how this licence infringement could actually do any significant damage to anyone — so while I acknowledge that this problem exist, I refuse to accept that it's as significant enough as people make it out to be.
For a long time, it seemed like technical progress in computing have stopped, and now that AIs and LLMs are finally bringing exciting new technology to life, it's very sad to see exactly the people that should be excited and inspired about it — software engineers — fighting against it.
Crypto bros are now AI bros huh?
Ideally, github should check the license of the code it's using to feed copilot, and only use code with permissive licenses.
I don't think this will fix anything.
Actually developers are the only ones that stand to lose here since now open source will be spread on multiple platforms, making it harder to find what you want
At the very least, they can't claim that you gave them special rights through the terms of service.
Also, rate limiting.
People want the benefits of publicly sharing stuff, but then they want to prohibit others from learning from what they share.
There are many options to keep things private. The downside is that you won't get the same exposure.
Conversely, open source licenses explicitly state that an end-user may further distribute that source code to anywhere they wish.
A question would be if creating and training Copilot is "improving the Service over time". I would suspect that it would be, though.
There are still some open questions around what happens when Copilot suggests code verbatim, but these are mostly for the users of Copilot. Although I would hope that GitHub is thinking about offering information to ensure that users understand the source of code they use, if it may be protected, and what licenses it may be offered under. There are still some interesting legal questions here, but I don't think that the training of Copilot is one of them.
A more interesting question would be what GitHub does if someone uploads someone else's copyright-protected code to GitHub and it is used for training Copilot before it is removed. If you don't own the copyright, you can't grant GitHub the rights needed to use that code for anything, including improving the service.
Definitely an interesting case to be had, but I'd argue that it does not. They're using their customers' code to create an entirely new product that would not be possible without it, not just improving their ability to host a Git repo. Otherwise, what standard is beyond "improving the service over time?" Can they do anything with the code they host as long as it improves their service? What about sell bootleg copies of it and use the proceeds to upgrade their servers?
The linked-to document explicitly DOES NOT prohibit others from learning what they share.
Quoting it: "If the project is under an open source license, it means that everyone can share a copy – even on GitHub – of the licensed material under certain conditions. A license restricting this right wouldn’t be open source anymore. However, since GitHub may not respect the terms of licensed code that is hosted on their servers, not uploading the code of others there is, in fact, an ethical choice."
Ingesting from public sources outside of GitHub will just become more necessary as they work to improve these things.
1. https://www.theguardian.com/commentisfree/2023/may/08/ai-mac...
I don't know if it's laughable, or utterly terrifying giving such a powerful tool into the hands of reckless wannabe entrepreneurs.
And something tells me that even if CoPilot would be entirely prevented from doing that, they would still not be happy about CoPilot using their code for training. The copyright issue is just a convenient pretext.
Although I don't use github (for reasons unrelated to to copilot), public access to my code was eliminated when I took my websites down while I look for a way to deal with AI scraping. I'm eagerly watching what others do, hoping that someone will have a great idea of how to deal with this before I do.
Remember the old saying: Anything published on the internet stays on the internet.
What prohibits someone from crawling code from other sites and building a GitHub Copilot equivalent?
Considering how ChatGPT style bots are often trained on public websites that is likely to already be true even.
With neural networks it's impossible to actually describe what the network does. Including proving that it has or has not used GPLed code for a certain input/output set.
One could argue that all code output is GPL with the associated restrictions unless said network has provably never been trained on GPLed code.
But I don't know how much success the author will have, in his endeavor.
The horses have fled. Closing the barn doors, does nothing. The Rubicon has been crossed. The die is cast. Iacta alea est, etc.
If you don't want your code public, don't make it so. Because licenses are a thing of trust and you shouldn't trust anybody.
*.code-workspace
.github/
.vscode/
in your .gitignore is another way to express that you don’t want Microsoft as a part of your project.According to the Hitchhikers Guide to the Galaxy…
A new coding highway was planned years ago. You could have filed a complaint in the Implications of Future AI Department in the basement of Bill Gates’ mansion, but you didn’t.
The highway is being constructed and you’re just going to have to deal with it.
Here’s some Vogon poetry to make you feel worse and a towel to cry in (only one of its many uses).
The guide also suggests gardening as a way to reduce anxiety.
Nobody can stop this anymore. If it's not github that's only taking code from their own platform, there will come others that scrape the internet and use everything, just like MidJourney. Embrace it.
I'm more against Art AI generators, because they produce end results and will fully replace the need to know how to draw. Copilot is just providing simple snippets, it's not thinking for you, so it won't replace the engineers completely, at least for now.
It's a better solution than don't use GitHub at all. And whether we like it or not, a new license should be introduced to address the new data model.
Yes programming is going away, so is most intellectual and artistics tasks.
I think github, after stackoverflow, is the best thing that happened to developers.
> I think github, after stackoverflow, is the best thing that happened to developers.
I don't think github is bad tooling (there are worse), but the best thing for developer? I don't see anything on github that is greatly better than what we have in other services.