2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.
2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.
Unfortunately the original report (https://www.theinformation.com/articles/alphabets-google-and...) is hardwalled.
Amplification of biases, propagation of errors, echolalia and over-optimization, lack of diverse data, overfitting
And your question about how OpenAI prevents their training data from being corrupted is one we should be asking as well!
Yes, this is an existential problem for Google and training future LLMs.
See also, https://www.theverge.com/23642073/best-printer-2023-brother-... and https://searchengineland.com/verge-best-printer-2023-394709
... it's uncanny how it always finds what you thought you were looking for!
Seems trivial. Only use old data for the bulk? Feed some new data carefully curated?
I'm only half joking.... I think we likely will end up with flags for human generated/curated content (and it will have to be that way round, as I can't imagine spammers bothering to put flags on AI-generated stuff), and we probably already should have an equivalent of robots.txt protocol that allows users to specify which parts of their website they would and wouldn't like used in the training of LLMs.
It will need to be based somewhat on the honor system (just because someone's proved they're a human doesn't mean they won't put their attestation on auto-generated text), but it definitely sounds better than nothing.
They'll still need to incentivize it somehow, though. Why do I as a human want to add that meta tag? If the answer is "better search ranking" then it renders the whole scheme mostly pointless because obviously spammers will want to acquire the attestation and attach it to their auto-generated content.
before: organic (south america) and regular (central ou SEA) for 69, 59.
then: both chikita's brand with regular and organic stickers (clearly the same produce, always from SEA) for 49 and 39 cents.
thats was days after the announcement
[0] https://www.amazon.com/Fumbling-Future-Invented-Personal-Com...
Apple wasn't the first one to try to make a successful smartphone, but they had resources, know-how, and tried at a better time with fewer unknowns around.
Google Maps wasn't the first map product. We all used mapquest way before. But Google Maps was technologically advanced. Ajax made maps usable for the first time.
Gmail wasn't the first webmail. Hotmail had millions of customers already. But Google gave people unlimited space to store old email, whereas email in the old days filled up your inboxes and needed to be deleted.
Question is if they can and will leapfrog.
Google Plus was a sign of desperation and utterly failed. The dozens and dozens of different Messengers (sorry I don't even know what the latest one they're pushing is, RCS?) all failed.
As an organization we will see in the coming months and years if Google can still overtake others when coming in from behind or not.
This isn't correct. Gmail launched with 1GB per user, which was way higher than other services, and they did keep doubling the storage space year-after-year, but it was never unlimited until Google Apps offered unlimited storage for businesses and schools.
In 2004 1GB of email was effectively unlimited though. Keep in mind hotmail offered 2MB and only increased that 125x when Gmail launched.
https://www.cnet.com/tech/services-and-software/hotmail-to-o...
https://arstechnica.com/information-technology/2011/02/googl...
A closer equivalent would be if someone had made a ShareSERP site and people posted their favorite search terms and the results Google gave and Bing crawled that and incorporated the search terms to links connections into their search graph.
The actual actions had maybe gone too far (personally I thought it was more funny than "copying"), the hypothetical would be pretty much what you'd expect to happen. Even google would probably crawl ShareSERP and inadvertently reinforce their own results (the same way OpenAI presumably gets more than a bit of their own results back at them in any new crawls of reddit, hn, etc even if they avoid sites like ShareGPT deliberately).
What now? Seriously?
I found this. Section D4.
"We need the legal right to do things like host Your Content, publish it, and share it. You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video."
"as necessary to provide the Service" seems critical.
> You retain ownership of and responsibility for Your Content.
and section D4 says:
> This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service, except that as part of the right to archive Your Content, GitHub may permit our partners to store and archive Your Content in public repositories in connection with the GitHub Arctic Code Vault and GitHub Archive Program.
There is nothing in the terms that requires the GitHub user to relinquish all licensing rights.
https://docs.github.com/en/site-policy/github-terms/github-t...
Under definitions: The “Service” refers to the applications, software, products, and services provided by GitHub, including any Beta Previews.
The terms make clear that uploading code to GitHub gives GitHub the right to "store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time" while the code is hosted on GitHub.
However, that's not the same thing as relinquishing (giving up) licensing rights to GitHub. The uploader still retains those rights, and there is nothing in the terms that says otherwise.
GitHub would argue that it is, and they'd likely argue that charging for access to copilot is akin to charging for access to private repositories.
Others would say that copilot is somehow separate from the services Github provides, so using their code for CoPilot wouldn't be covered by the ToS.
I'll repeat the definition of service: The “Service” refers to the applications, software, products, and services provided by GitHub, including any Beta Previews.
Fortunately HN commenters are not judges. And I would wager any bet that MS lawyers would not try to argue based on their ToS either, that would be a recipe for loosing any court case.
If you're posting anything you did not create yourself or do not own the rights to, you agree that you are responsible for any Content you post; that you will only submit Content that you have the right to post; and that you will fully comply with any third party licenses relating to Content you post.
I suppose this means if I upload your stuff to GitHub, and you sue GitHub, then GitHub would be able to somehow deflect liability onto me.
> You may convey verbatim copies of the Program's source code as you receive it, in any medium, provided that you conspicuously and appropriately publish on each copy an appropriate copyright notice; keep intact all notices stating that this License and any non-permissive terms added in accord with section 7 apply to the code; keep intact all notices of the absence of any warranty; and give all recipients a copy of this License along with the Program.
https://www.gnu.org/licenses/gpl-3.0.en.html
If GitHub then uses the source code in a way that violates the license, there is no provision in the GitHub terms of service that would allow GitHub to deflect legal liability to the GitHub user who uploaded the program. The uploader satisfied the requirements of GPLv3, and GitHub would be the only party in violation.
If you can't actually grant that separate license, you're misrepresenting your ownership and license to that code
> If you upload Content that already comes with a license granting GitHub the permissions we need to run our Service, no additional license is required.
https://docs.github.com/en/site-policy/github-terms/github-t...
and section D4 does not mention any permissions that GPLv3 does not already cover. GitHub automatically recognizes when a repo is GPLv3-licensed, so it cannot claim ignorance of what GPLv3 is.
OpenAI and Google both scour the web for human-generated content. What Google cares about here is the learnings from OpenAI's proprietary RLHF dataset, for which they had to contract a large sum of human labelers. Finding a roundabout way to extract the value of a direct competitor's purpose-built, costly data feels meaningfully different from scraping the web in general as an input to a transformative use.
OpenAI and Google both scour the web for content, period. That content could be human generated or AI generated or a mix of the two. Neither company is respecting copyright or terms of service of every individual bit of data collected. Neither company cares how much effort was put into creating the data, whether humans were paid to do it, or whatever else. So there really isn't that much difference between the two. In fact I can guarantee that there was some Google-generated content within OpenAI's training data.
It's like the guy who never brings anything to the potluck, but after everyone finishes eating, he boxes up the leftovers, and starts selling them out of a food cart.
That particular example doesn't seem all that great, as it gives the impression of not letting food go to waste.
Though sure, the guy could just give the food away, but shrug.
Yes, this latest instance with OpenAI outputs is shady, but I think it's in the same spirit as scraping news organizations for content which journalists were paid to write, and then showing portions of it directly in response to queries so people don't go directly to the news organization's pages, and it's in the same spirit as showing answers to query-questions that are excerpts from scraped pages which another organization paid to produce.
There we go again, its, one law for the unwashed plebs and the other for us.
Why do you think that I, after spending my time and effort to write my blog, own my content to a lesser extent that OpenAI does their? Such hypocracy.
I really don’t understand this angle. In fact, I am fairly positive that the training set for GPT-4 contains many thousands of conversations with AI agents not developed by OpenAI.
Do AI companies need to manually sift through the corpus and scrub webpages that contain competitor LLM output?
(“Yes” is an acceptable answer to this, but then it applies to OpenAI’s currently existing models just as much as to Bard)
Not legally the same situation, but ethically close enough.
Read their statement carefully and it's actually not a denial of the allegation.
> But Google is firmly and clearly denying the data was used: “Bard is not trained on any data from ShareGPT or ChatGPT,” spokesperson Chris Pappas tells The Verge
* Allegation: Google used ShareGPT to train Bard.
* Rebuttal: The current production version of Bard is not trained on ShareGPT data
Both things can be true:
* Google did use ShareGPT to train Bard
* Bard is not currently trained on any data from ShareGPT or ChatGPT.
It depends on what the meaning of is is ;)
Did they accidentally train on that public piece of info they scraped anyway because they are scraping the whole web?
Or did they intentionally scrape chatgpt output to see if that would help?
Then after, train on raw data.
This association makes no sense.
If the latter, there are many laws that say you can own an idea, provided it exists somewhere.
But the big news is that it works, just a bit of data can have a large impact on the open source LLMs. OpenAI can't have a moat in their proprietary RLHF dataset. Public models leak, they can be distilled.
The idea of doing this is embarrassing enough for Google.
Google index the whole web, some of the documents are due to be generated by ChatGPT, there is no way around it.
My stronger opinion is that the people who can do this stuff via having a crawled corpus of the Internet need to keep in mind that it's all our "user-generated content" that they've freely appropriated to build their models, and so whatever the technical copyright rules are (or become): you don't ethically own something that's closely imitating stuff we all wrote over the years.
I think the argument here is over the OpenAI Terms of Service, not copyright.
Seems to me that’s an issue between you and OpenAI. (Does your blog or code repository actually have published restrictive terms of service? Did it when OpenAI accessed it? Did OpenAI even access it?)
Microsoft is out there laundering GPL code with Copilot. These companies live firmly in the don't give a fuck region of capitalism. Copyright law for thee, not for me.
Maybe if they had put in their terms of service "you can only share this on sites with their own ToS that allow sharing but disallow using the content for training models, and also replicate this requirement", I don't see how you could have any sort of viral ToS like that.
Seems more like it's just a bad idea to rely heavily on another LLM's output for training.
but it's google so no big deal right?