Generative AI profits off your code. Make them pay for it
paytotrain.ai
paytotrain.ai
> We need the legal right to do things like host Your Content, publish it, and share it. You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video.
Doesn't that contradict the purpose of this website? Is this performance art?
Also, why is this website so secretive? Why not publish the license on the website?
> Who is PayToTrain created by and why?
> PayToTrain is created by a small group of developers and attorneys who are passionate about open source software and ensuring that developers are properly compensated for their work. The website and service are provided completely free of charge.
Edit: PayToTrain looks like a non-disclosed ad and/or project from legalist . com.
So if the clause is widely adopted, it may be good for Microsoft and bad for Salesforce. If you want to reward Microsoft and punish Salesforce, it may be a good idea.
If Microsoft loses this case, it actually means Microsoft wins and we all lose.
Who has a large enough corpora of training data? Only institutional copyright holders.
This is probably going to play out like Oracle vs. Google when Google suddenly realized that they should lose and intentionally threw the case.
I'm so worried about this case. Treating copyrighted training data as fair use, and letting models learn as a child might learn from a book or movie, is the best way to proceed. It widens the playing field for both development and disruption.
The thing is, said person both reads way less code than a non-biological neural network, and emits its derivations based on many inputs regardless of the code it ingested via its high resolution multi focal adaptive light sensors. Including but not limited to experiences, communication with other biological neural networks, human-machine code translators (compilers), daily unpredictable hormone fluctuations and infinite other inputs it processed which affected all aspects of its cranial muscle and daily living circumstances and choices.
IOW, A neural network is neither a person, nor learns the same way or derives and emits the same way.
This is equal with claiming that a Furby is a person, just because it can babble and blink.
In a similar way, noone really understands intuitively how these ML models are actually working (we treat them as mostly black boxes in practice), in contrast to looking at an equation for instance. I have played with some of these text generation models, and frankly we are already at the point where deciding whether or not they pass the turing test depends on the details rather than the spirit of the rules for the test. It may not be a coincidence that NNs designed to replicate our own brain structure also replicate important aspects of our cognition.
>Neural networks approximate the function represented by your data.
Meanwhile, the institutions will leap ahead of us. Models and annotated data sets will forever be out of reach. Open source equivalents will be severely behind the status quo.
An intentional open source dataset could target new domains that there was no institutional will to pursue. I strongly believe that the open source community's capabilities far exceed that of any single large corporation.
What are you talking about here?
Which also opens the floodgates for bleeding GPL and other copylefted code to proprietary realm. Very convenient, yes.
I’m never OK with someone making a derivation engine which offers my GPL code to a closed source base.
But Google won?
Add our “Humans Only Clause” to your MIT license. Your code is still open source — for human developers only.
Sore disappointed that there is no entertainment involved. That's actually a pretty cool idea.
So github doesn't have (could be wrong) a default license grant or a over-riding licensing agreement. Your project, your license. If you change the license of your project, that is entirely your choice.
As to the Q of should we be generous to our corporate masters or take this opportunity to stick to the man and get rewarded for our mind products and compensation for being geeks! Society does owe something, does it not? /G
It's worth having a discussion about it, imo.
Teaching AI how to code is a continuation of building this ecosystem because people will use these generative AI coding tools, lowering the bar to code.
You could argue that training is a form of compilation, and the weights are a derivative work.
In any case it raises some other questions about intellectual property as a whole. If you can sue an AI model for profiting off your intellectual property, why can't you sue a human for the same? Say you read a book one day, and are so inspired that you go ahead and write a new book. Imagine you publish that new book and sell millions of copies. Are you entitled to pay royalties to the author who gave you inspiration? It seems to me that unless you're plagiarizing large chunks of the original work verbatim, you probably shouldn't be forced to owe the original author much of anything. LLMs do plagiarize, but they do so somewhat inconsistently due to their non-deterministic output (just like humans!).
So in this case the weights are not even a derivative work but just a compressed copy of the original code.
However, have they reviewed Microsoft's claims that their use of code for Copilot is open source? And if they have, is there somewhere I can read that analysis?
Now, Microsoft could be wrong on that claim, but until someone convinces me otherwise I'm going to assume their lawyers did their due diligence and they're correct. If Microsoft is correct about that it doesn't matter what you put in your license, and thus this is useless.
"their use of code for Copilot is open source" -> "their use of code for Copilot is __fair use__"
> PayToTrain is created by a small group of developers and attorneys who are passionate about open source software and ensuring that developers are properly compensated for their work. The website and service are provided completely free of charge.
That doesn't answer the obvious implied questions of how much anyone should trust this effort -- such as not to be selling them out to the infingers, either individually, or as as a "self-regulation" model that the infringers can point to in upcoming legal battles. And I'd be surprised if the attorneys didn't realize that.
Also, some of my open source projects have an explicit clause which states that 'The source of this code shall not be misrepresented' so if chunks of my code showed up in some AI-generated code without attribution, it would be a pretty clear-cut case.
https://twitter.com/docsparse/status/1581461734665367554?lan...
In this instance the weights of the AI system just contain an obfuscated copy of the original source code.
I am not a developer, but I understand how generative AI leveraging your work to make it easier for someone else.
A similar thing needs to be done for images too.
To me this sounds like it's antithetical to open source software because the point of making software open source is so that other people can leverage your work. It shouldn't matter if it's done through generative AI or through a human's brain.
The point is that other people can leverage your work under the terms you distribute them under. For the vast majority of open source licenses, that means giving attribution and including the copyright notice and license when distributing the source code or its derivatives. For others, it means all of that and releasing derivatives under the same license.
If developers wanted to distribute their code under licenses with different terms, they would have, but they didn't.
People often write code that looks like existing code that they've seen even if they're not aware of it, it's a blurry line. I see it as just banning AI from doing the same thing as humans just because it's better at it.
An argument could be made that it's fair for an AI to not attribute the code it outputs too. The human-human reason for attribution is "I wrote this code by doing X amount of work, since you're using it and it'll save you time, I should fairly be given attribution". But then the AI is also writing out the code that it's prompted for, it's just faster at doing it.
Why not create a tool that instead runs the AI generated output through a check that provides proper attribution? Then you'll also get human written code that doesn't attribute the original author as well.
I would be fine with an AI being trained on my code, provided that the weights of said AI would then be published under the same license as my code.
Is there a license for that?
A legal grey area exists as to whether publicly available creations (code or art) can be used to train datasets for generative AI projects without infringing their creators' underlying copyrights. Other types of claims, such as violation of license agreements and DMCA violations, require proof of damages to substantiate.
The legal solution we’ve identified is to add a specific damages amount to the license itself — a licensing fee. The failure to pay such a fee would cause the creator to suffer damages in the amount of the fee. By imbedding a licensing fee into a traditional open-source license, a creator can solve the proof-of-damages issue that could otherwise limit a claim under the DMCA or for breach of contract, and limit the fee to generative AI companies.
That’s why we built the Humans Only Clause. If you don’t want your code used by Copilot in this way, the Humans Only Clause can help strengthen your protections from use for training purposes. It’s a simple addition to your existing open source license to keep it free use and open source for other developers, but to prevent use without attribution by generative AI companies.
You can access the Humans Only Clause and insert it into your GitHub repo by going to PayToTrain.ai — we also built a payments form where you can set your own licensing fee depending on how valuable you believe your repo to be. If we get enough people using this clause, there’s a good chance we can assemble a separate class for a future class action, where each user gets significantly higher damages than what’s available statutorily under existing DMCA lawsuits.
On a philosophical level, we believe that the open source community is based on principles of taking and giving back to the collective. AI-based programming assistants strip away any attribution while drawing from the underlying contributions of the community. We want the open source community to continue to be open source, but we don’t want big companies to profit on our code.
If you’re interested, check it out: paytotrain.ai. We’d love to hear what you think
How is this behavior different than 90% of human coders? Most SW devs scream if you ask them to pay for something, whether its apps, code, TV, movies, etc. But they will happily try to build startups on top of a vast mountain of free code. I really don't care if my code gets sucked up by the AI vacuum, humans have been doing that quite well for awhile now.
Here was my attempt to write a clause prohibiting language model training/inference:
https://bugfix-66.com/7a82559a13b39c7fa404320c14f47ce0c304fa...
3. Use in source or binary form for the construction or operation
of predictive software generation systems is prohibited.
How does the Humans Only Clause fix the flaws in my attempt?The Humans Only Clause adds an explicit licensing fee, and what else?
How is the clause worded?
> The license must not restrict anyone from making use of the program in a specific field of endeavor. For example, it may not restrict the program from being used in a business, or from being used for genetic research.
It also seems to violate freedom 0 of the FSF's four essential freedoms that define free software:
> The freedom to run the program as you wish, for any purpose (freedom 0).
I'm not sure this can be used by open source projects if they want to remain open source projects.
If that confused you, or you consider it "obnoxious", then you are not the target audience.
That's ok. Hacker News is not 100% hackers!