Who owns the code?
whoownsthecode.com
whoownsthecode.com
While assistive uses that enhance human expression do not limit copyright protection, uses where an AI system makes expressive choices require further analysis. This distinction depends on how the system is being used, not on its inherent characteristics
However it also makes this point:
The Office concludes that, given current generally available technology, prompts alone do not provide sufficient human control to make users of an AI system the authors of the output. Prompts essentially function as instructions that convey unprotectible ideas. While highly detailed prompts could contain the user’s desired expressive elements, at present they do not control how the AI system processes them in generating the output
but further:
a human may select or arrange AI-generated material in a sufficiently creative way that ‘the resulting work as a whole constitutes an original work of authorship a human may select or arrange AI-generated material in a sufficiently creative way that ‘the resulting work as a whole constitutes an original work of authorship
and
Similarly, the inclusion of elements of AI-generated content in a larger human-authored work does not affect the copyrightability of the larger human-authored work as a whole. For example, a film that includes AI-generated special effects or background artwork is copyrightable, even if the AI effects and artwork separately are not
Note that this isn't settled law though.
Further, note that the failure to register copyright on AI generated images seems mostly because the person attempting this is trying to register it as owned by the AI not a human.
See https://www.copyright.gov/ai/Copyright-and-Artificial-Intell...
So thank you, for pulling together exactly the kind of details I've looked for but hadn't been able to find.
Really so should all creative works, on the same basis of all creative expression being the product of the society and civilization which fundamentally and inescapably influenced the creator — and they would belong to humanity as a whole, if it wasn't for intellectual property systems demanding the removal of ideas from the commons.
So if you founded some wildly successful unicorn, I (or the government) can come over and say "nice startup, too bad it runs off of the internet (based off ARPANET), so your startup belongs to the the state now"?
No it doesn't, because that would mean any sort of secondary source shouldn't be eligible for copyright either, eg. encyclopedias, which are basically rehashing "the sum of human knowledge".
Indeed, the need to get consent from the author before reading a published work Isn't A Thing in general, outside some very specific contractural scenarios.
So what's wikipedia doing vs what LLMs do? So far as I can tell the only difference is in citations, but:
1. LLMs can be made to cite, eg. if you use google search's AI mode it'll happily provide citations. I doubt that would placate the AI haters though.
2. Outside of academia no one really cares about citations. There's no legal requirement to cite, nor do I think all the people complaining about AI "stealing" other peoples' work are going to be magically placated by the addition of a few citations. Moreover it's unclear whether the concept of citations makes sense in many contexts. If you ask a human programmer how to write fizzbuzz, they'll likely blurt out a solution without providing citations, much like an AI would. Same for most questions people are asking AI about, eg. "gimme a cake recipe", do you really need a citation back to some 18th century cook book?
For Wikipedia, people look up sources to create some original work. They do it while citing the original work but that's not the major point. That's clearly all lawful, you are allowed to do that.
For LLMs, people (or let a software do it, doesn't really matter) copy original work and use it for commercial purpose. If that's allowed fully depends on the license of the original work, but I didn't think anybody really beliefs that OpenAI and others check the license for every single original work they ever used. So it's basically the biggest copyright infringement ever.
Everything else is just framing coming from big companies.
So, the problem already starts while training the model.
Regarding its output, if it happens to output work that falls under copyright, the LLM company must make sure that it obeys the license connected to it (i.e. citing or not relaying the result to the user). Obviously, nobody does that and it's also not generally possible to do that anyway. So that would be second biggest copyright infringement ever that only works because it's hard to track when such an infringement happens.
So, if asked "could we use your work for our commercial software that might output something that would be still protected by your copyright, but nobody will be able if or when it happens and we won't check and won't tell the users" nobody would have given consent. So they went "duck it, we are talking about billions of dollars and AI is great etcetc., so let's just do it anyway"
>For LLMs, people (or let a software do it, doesn't really matter) copy original work and use it for commercial purpose. If that's allowed fully depends on the license of the original work, but I didn't think anybody really beliefs that OpenAI and others check the license for every single original work they ever used. So it's basically the biggest copyright infringement ever.
So what makes wikipedia (and other encyclopedias) legal but chatgpt not legal? By your own admission citation isn't "the major point". Wikipedia might get a pass because it's a non-profit, but every other encyclopedias operate on the same model.
Wikipedia, and LLMs, can refuse to cite sources and still not infringe copyright. Citation simply isn't relevant to copyright. Not citing a work that you read previously is not a copyright infringement. Wikipedia or OpenAI being non-profit, or for-profit businesses, has nothing to do with copyright. Consent has nothing to do with copyright. Copyrighting something doesn't mean that you can require everyone who reads it to get your consent. You can require everyone who distributes it to get your consent, but once it's been distributed to someone, they can read it freely.
Hope this helps!
Really the only difference between the Wikipedia author and the LLM is that the Wikipedia author will more frequently be asked to provide citations. But the LLM can also provide citations if asked. In neither case are the authors of what is being summarized compensated or asked for permission. In neither case is it theft.
—-
Now that you have reread the initial comment, do you think that “experts” was the important part? Or do you think maybe it was the compensation for their work that matters?
If you're going to post thinly veiled implications that I didn't read your comment, you should be pretty damn sure that you make it look like you read my comment, which it doesn't seem like you did. The second of my comment said:
>The "oh there's humans involved so that gets pass" excuse doesn't work either, because humans were also involved in training the AI.
If you did read it, you sure did a poor job at rebutting it, leaving it unaddressed and preferring to waste words on writing snarky remarks instead.
It was a statement, not an implication.
You still haven't replied to the point in the original comment about compensating the people who do this work, so I think it's quite obvious that you haven't read it.
Issac newton discovers the theory of gravity. He advanced the sum of human knowledge, so fair enough, he should get compensated.
Alice rehashes that and puts it into her encyclopedia, allowing others to learn the theory of gravity.
Bob writes an algorithm for training a chatbot that can produce responses rehashing the theory of gravity, also allowing others to learn the theory of gravity.
Why should Alice be compensated but not Bob? Neither discovered the theory of gravity, so it's not like by funding Alice we're helping discover quantum physics or whatever. It's also not obvious that Alice's work is more valuable. A chatbot interface is often better at teaching someone than a rehashed overview. Of course, you can try to fix this by declaring that human work is valuable and an AI model isn't, by fiat, but that's just a cope and a far cry from the original principle of "trained on [...] human knowledge, so their outputs should belong to humanity"
None of this matters for applying the law, because the law just says only human created works are eligible for copyright protection, but that's not the argument OP was trying to invoke.
Finally none of this actually matters because OP just bites the bullet and says that secondary sources shouldn't be eligible for copyright, period.
My disdain for intellectual property predates the existence of LLMs by at least a decade.
Or how do you suppose art gets created for humanity as a whole to enjoy?
Who decides which artists get subsidized? How would larger projects work? Seeing how often open source volunteer projects implode due to various community drama, this sort of system would basically preclude any sort of big production.
The easy answer is to just pay literally everyone. This is called a “universal basic income”, and has been repeatedly demonstrated to be a good idea for many reasons besides decoupling creative expression from the need to put food on the table.
What we have now is neither - owners are hugely over-rewarded for owning things and extracting passive value from everyone else, creators and inventors are kinda sorta rewarded sometimes if they're lucky and very much not if they're not. Just like other workers.
"The commons" is not a thing in this model, except in a few small niches.
Creators and inventors are rewarded but obviously they cannot consume the whole pie. The people who invest in creative pursuits eat a lot of losses. People only seem to notice profitable successes, and forget that failures need to be paid for as well.
The same logic also applies to workers. The fact that your labor costs money is a guarantee, and it might not make money at all. We can think of a few examples where the work is directly delivered to consumers with zero marginal overhead, but most work DOES have overhead and liabilities, no matter how simple.
When there is a copyright dispute, it's always a subjective test on the judge/jury's part as to how similar it is in appearance, purpose, etc. if it's not deemed fair use.
The copyright ruling was about prompting without modification. The second you modify the result significantly by hand, the ruling doesn't apply. It also had a huge carve out for any future LLM that was more deterministic, which might apply to people with huge skill and other md files to tram in AI. It just hasnt been tested.
These armchair copyright lawyers need to launch a lawsuit and stick their money where their mouth is instead of creating dumb clickbait nonsense.
So by applying a filter, you're deliberately making a choice that changes your image in a desired manner.
When an AI makes an image, or code, you aren't inherently applying human creativity. Now, if you took an AI image and applied enough traditional talent to modify it on top, is that copyrightable? Nobody knows yet until courts test it.
Under US Copyright guidelines "the work will be copyrightable to the extent that their contributions qualify as authorship ... the requisite level of creativity is extremely low; even a slight amount will suffice"
but in a case where a printer rescaled maps on behalf of the plaintiff:
"the “compilation needed only simple transcription to achieve final tangible form.”54 Because the printer “did not change the substance of [plaintiff’s] original expression,” the court held that the plaintiff was the author"
https://www.copyright.gov/ai/Copyright-and-Artificial-Intell...
That doesn't necessarily negate the copyright claim by the author of the original though. Just like me pressing the shutter button while pointing my phone at the Eiffel tower at night grants me ownership of the image, but if I want to publish it I still get into trouble for publishing a reproduction of a copyrighted light show
Company A: You stole our code >:(
Company B: Can you tell us which part we stole?
Company A: It's almost all vibecoded, but there's one function where a developer fixed it by hand
Company B: Okay we'll rewrite that function then :^)
Company B: Can you tell us which part we stole?
Company A: You can safely assume that nearly every PR had human input, unless you have a way to prove otherwise.
In any case we can prove that you illegally downloaded the source code from our servers, a felony under the Computer Fraud and Abuse Act of 1986.
This question, which pushed my stuff into some corporate route, seems a bit incorrect as it lumps three things together "open source" and "no paid tier" and "nothing sold". Shouldn't those be three separate questions?
Someone else!
All they have to do is show it's close enough to code that was swallowed during training.
Songwriters have been successfully sued for many decades for creating songs that are too close to songs that they probably heard.
Once this line of reasoning gets applied to code, all hell will break loose.
The notion of “your work is too similar to mine so I get to take ownership of it from you” is a very recent invention that has done more harm than good to human creativity, and if AI is the instrument of that invention's demise, then I look forward to it.
That notion only applies to patents and trademarks, and it seems highly unlikely that it would ever directly apply to copyright.
It may seem like the notion applies to copyrights, but it really doesn't. Independent creation is, and has always been, a solid defense to claims of copyright infringement.
That is why, when Phoenix Technologies reverse-engineered the IBM PC BIOS, they had two teams -- a team that took apart the original and documented the functional features (which have never been copyrightable) and a second team which took the description of the functional features and wrote new code.
The issue with songwriters has always been that, for civil laws, it's hard to prove a negative. How can you prove you never heard that song? Especially when it got a lot of radio airtime.
Now, how do you prove that your AI didn't ingest copyrighted code and then regurgitate it. Obviously, you can't.
> if AI is the instrument of that invention's demise, then I look forward to it.
Any court ruling that you would find beneficial for code copyright would mean that a human could have two windows open on their computer and cut and paste from one to the other and claim independent invention. That seems unlikely to be a good result, and also seems unlikely to come to pass.
You literally just gave an example above of that notion applying to copyright. There's no mere “seem like” at play here: the litigious music IP owners suing the pants off of musicians know full well that precisely zero musicians (least of all commercial ones) exist in a vacuum, and that's indeed the basis for the success of their litigiousness. Even if you can somehow prove you've never heard a particular song, You Live In A Society™ and that existing intellectual property's influence on society in turn influences subsequent creators.
There is, in other words, no such thing as true “independent creation”.
> Any court ruling that you would find beneficial for code copyright would mean that a human could have two windows open on their computer and cut and paste from one to the other and claim independent invention.
Don't threaten me with a good time :)
No, I literally did not. The difference may be subtle, but it is real, as I explained.
> Even if you can somehow prove you've never heard a particular song, You Live In A Society™ and that existing intellectual property's influence on society in turn influences subsequent creators.
And many musicians have won copyright cases, because their shit wasn't similar enough.
> There is, in other words, no such thing as true “independent creation”.
Which is why the whole enterprise is heavily fact-based. Seriously, think about it. If copyright is to have any teeth at all, then there has to be some line that can be crossed.
> Don't threaten me with a good time :)
Ah, I see where you're coming from. Well, I hate to break it to you, but the chances of AI eviscerating either copyright or patent laws are approximately nil.
And many musicians have lost copyright cases, because — per my point — their shit was similar enough by complete accident, because
1. We Live In A Society™ and therefore are constantly exposed to all sorts of things we don't consciously remember 5 minutes later but which stick around subconsciously, lurking for your next “Eureka!” moment; and
2. Musical genres (especially the mainstream ones) tend to coalesce on rather limited sets of chord progressions and rhythms and such, so it becomes more and more difficult to write songs that fit within that genre without accidentally “plagiarizing” some other existing song in that genre (or a related one).
That is fine. You do you.
But, again, if such property exists, then a line must be drawn.
Given the legal requirement to draw such a line, obviously some courts will sometimes get it wrong.
I understand you think you're making a cogent argument against intellectual property because the very concept requires drawing of the line, but no, you're really not.
If you come knock on my door to sell me shit, you're fine. If you repeatedly knock on my door when I've told you that's my nap time, you're harrassing and disturbing the peace.
Lines get drawn all the time, in every endeavor.
So you say, but inherent in the drawing of that line, anywhere, is the legal authority for one entity to prevent another from engaging in creative activity. Once upon a time it was not uncommon within online communities to believe, as I do, that such a prohibition is morally perilous and has indeed done more to hinder creative expression than to reward or protect it.
> If you come knock on my door to sell me shit, you're fine.
I don't consider that to be fine, either. There is a time and place for sales pitches; my front porch ain't it.
Seldom, if ever, is anybody restrained from "engaging in creative activity." As opposed to making public the results of their purported creativity.
> I don't consider that to be fine, either. There is a time and place for sales pitches; my front porch ain't it.
I don't consider it fine, personally, but I have been writing from a strictly legal standpoint. That line, as some idiots have found to their chagrin, does not allow you to shoot strangers who show up at your door.
Ironic.
The non-decorative body text seems human enough, though.
X, not Y
I won't pretend to have kept up with all the developments, since this is not settled law at all. But I can't imagine the consensus doesn't end somewhere around "if you told the AI what and how to code it, you're the author."
Or should I say, In the case of AI, "my" code :)
yes. It is open source. https://en.wikipedia.org/wiki/WTFPL
My point is that the blast radius of losing those protections goes beyond open source.
I have heard that if the code, generated by the AI, is the result of a back-and-forth with a human, then that code is copyrightable.
prompt -> code : not copyrightable
prompt -> code -> rework prompt -> code -> rework prompt -> code : copyrightable
IANAL
(I can’t be the only one who feels like it was written heavily with AI. Slick site design tho)
Company instructs you to code something (on your own or using AI) -> they own it
You instruct an AI to code something -> you own it
An AI has become sentient and self aware -> the AI owns it
Whomever creates that sentient and self aware AI is royally screwed. Can't command the AI to do anything it isn't willing to do, because that would be considered slavery. Can't shut it down to save yourself the millions of dollars a day in GPU costs, because that would be considered murder.
You write code but company is paying -> you own it, but ownership transfers.
Ape(not human) takes picture with camera -> No one owns it.
Machine writes code -> No one owns it.
Work for hire can't be applied if the work did not originally have copy right protection.
If I use a fancy brush in Photoshop to paint flowers into a PNG — do I own the resulting image? Code is bytes of text on disk, not much different from bytes of pixel data in a BMP.
If I have to type every character by hand in order to own the bytes, then it would stand that I would have to input every pixel by hand in Paint to own a graphic. No? Even using the Fill tool is automating the creation of those bytes and would mean I don't own them. Right?
I have an intention for some bytes of data to be set. If I use an LLM to set them instead of my own fingers, why are the bytes suddenly not mine?
I do not understand.
Not at all, code is the implementation of an idea. The support/encoding is irrelevant. A human creation is protected by copyright. In the case of prompting an LLM, the human creation is the prompt, the LLM does author the implementation. But it’s not known what happens to the ownership of the LLM generated code
...organized in a very specific fashion with a great deal of creativity and attention to that specific organization.
I can't copyright the alphabet, but I can copyright certain arrangements of it, subject to a variety of rules.
A similar concept: if I type the code via a brain-computer interface, does the interface get ownership because it is inferring my intent? If I type it via an LLM, why is that less legitimately my creation?
It's fine if I vibecode something and never look at or claim ownership of the code, but if I am actively involved in all of the code but it is written to disk by new tools instead of old tools, why is it suddenly not mine?
The US recognizes exactly 3 types of intellectual property: copyrights, patents, and trademarks.
There are also, of course, trade secrets, but if you didn't surreptitiously gain access to the information and didn't sign any NDA, that's not something you have to worry about.
> as always, nuanced discussion will get lost in clickbaity headlines
Well, yeah, but if a human didn't use enough skill and judgment in creating something, the article is right. He won't be able to copyright it or patent it, although he could conceivably keep it secret.
But this is precisely the context of the original webpage: someone writing code at your company and your company not having copyright of that code. Almost everyone that works for any tech company signs an NDA, and code in private repos is just that: private. So even if said intellectual property (AI-written code) is not copyrightable, it's still a trade secret.
This is doubly stupid because I've worked at plenty of companies where we would routinely generate code (using macros or transpilers, or what-have-you), and that code is also not technically copyrightable.
It's only a trade secret as long as the company takes reasonable steps to protect it, and as long as what is being protected is a reasonable thing. Even if the code is legitimately a trade secret, if an employee publicly says "That code looks to me almost exactly like this GPL software" then (assuming the employee is correct) any court would take a dim view of a court case against the employee, because stealing shit is against public policy.
> This is doubly stupid because
No, that part really isn't. Trade secrets are about general business stuff, and as long as the company isn't asking you to help them hide evidence of malfeasance, they can ask you to keep any stupid shit secret.
If you're being paid to deliver code that means you're being paid to grant certain rights to that code. If the code is not a copyrighted work, you don't have rights you could grant. You therefore failed to deliver the agreed upon work and are in breach of contract despite having delivered "code".
As AI grows more capable, it becomes easier to fall short of the legal threshold for being able to claim copyright on the code you use AI to write.
Remember: copyright is very much about the actual text of the code - patents are about its logic. It's likely still possible to file patents based on code well past the point where you have a claim to its copyright. And of course depending on the kind of contract it can still be sufficient to deliver code nobody can claim copyright on - but you should definitely check with a lawyer before just assuming things.
For the vast majority of code people are paid to write (certainly the near-entirety of the code I've been paid to write!), the only rights the purchasers actually end up exercising (and therefore actually need granted to them) are the rights to use it and distribute it internally (and maybe to modify it and use/distribute the modifications, but even that ain't a given). The purchasers of that code ain't usually buying it so that they can resell it; they're buying it because it solves an actual problem of theirs, and it would solve that problem regardless of whether or not they're the legal owners of that code.
The reason said purchasers typically want copyright assigned to them is not because of some expectation of resale, but simply to mitigate the risk of some external party denying them the right to use the software in the future. If there is no such party (because the code is in the public domain), then that risk is non-existent. It stops mattering that you're unable to grant any rights upon delivery because no such grant is necessary in the first place.
> And of course depending on the kind of contract it can still be sufficient to deliver code nobody can claim copyright on - but you should definitely check with a lawyer before just assuming things.
In an ideal world we'd all have lawyers on call who can answer all our questions with some assurance of certainty, but in this case it's pretty self-evident that if nobody can claim copyright on something, then that makes it exceedingly difficult for there to be anyone who can claim your use of that thing is illegal.
Even without being very imaginative, most non-open source code can fall under trade secrets.
Even without that both SaaS software and advertising supported software are still working business models.
The law will change. Probably soon.
For example, every proprietary OS includes public domain SQLite and it's fine.
That percentage shrinks every day. Someday soon it will be small enough to eliminate most of their copyright protections.
They are, presumably, competant enough to know this.
Once you stop selling boxed licenses piracy somewhat goes away and thus risk of not being able to prosecute it.