Stack Overflow simply bans folks who don't want their advice used to train AI
theregister.com
theregister.com
Lots more discussion:
Its not just people who are upset getting banned for being upset, its people who are attempting to burn it all down on their way out the door in protest.
I get where the protests are coming from, but the cardinal rule of online communities is "once you post it, its out there". Other people have reacted to it, replied to it, quoted it. You break not just your own content but entire discussions if you mass-delete your contributions. They stopped being exclusively yours to take back once you contributed them to a broader conversation.
Imagine e.g. the effect on democracy of politicians being able to have all of their past utterances deleted from the public record because they don't feel like it represents their current branding anymore.
For private individuals not "of public interest", the situation is a bit different, but e.g. the EU's "right to be forgotten" is arguably not a basic human right, but a relatively new and narrow concept for which there isn't any type of international consensus yet.
That's you opinion, not a cardinal rule. For example, Reddit comments and posts are deleted by users all the time.Same for Twitter, where users delete their posts and reacts all the time.
When I say "rule", I'm speaking in terms of "how each person posting content should operate". Once you make your words public, you (should) have no expectation of being able to take them back.
It doesn't let you delete your question after someone has gone to the trouble of answering though. Good thing IMO.
This asker ignored the rights of the answerers, moderators and SO, all of which put effort into creating that. I won't mind if someoen deletes his own unanswered question, but vandalising it after answerers put in an effort to answer? That's rude and egoistic.
If there is a delete button, it's fair game to use.
Yes, the lack of that changes how people participag on SO. I think it's a considerable reason why SO is the large resource that it is, with as many answers as it has. If askers could delete their questions (and thereby had the ability to make my answers un-google-able), I'm not sure I'd bother to answer, for example. Would you?
More to your point, it’s focused on evergreen information. Ideally the questions and answers get refined over time to resolve in to accurate documents that the public can use to solve problems. Most social media is about attention/dissemination in the moment and puts the individual above all.
I thought reddit keeps them still in the backend.
This only exists as a wish, but is not written in stone and some people actually disagree to that.
Whether it applies to Stack Overflow posts is another discussion. IMO it doesn't due to Creative Commons licensing.
As with almost anything GDPR: it really depends on the content, on the website, on the situation. There's no clear cut answer that applies to all cases.
And copyright law also enters into the picture here, which is why I added the second paragraph.
Specifically since what you've posted on SO you've shared as CC BY-SA.
Try publishing a book with any European publisher and then "right to be forgottening" it out of existence.
Says who, exactly? Stack Overflow, previously, did not say this. Reddit threads get deleted all the time. Facebook posts slide into the ephemeral "past" and are never seen again. Old forums die entirely, with nothing but (hopefully) a copy on the wayback machine.
And that's just regular old rot, in this case, Stack Overflow is explicitly changing the rules on what they can do with writing provided by volunteers under a previous understood agreement, to a new agreement, with no opportunity for negotiation and no engagement with their community. Simply an edict, issued from On High: "We can now sell your contributions to AI companies." And I'm sure for plenty of people they don't give a shit, but clearly some do, and so we get what HN is constantly clamoring for, individuals making decisions about their own creations and their own speech, but now it's suddenly bad because it's the Wrong decision.
Edit: source https://stackoverflow.com/help/licensing CC BY-SA
The point is: people don't want their code used to train code bots. Their wishes should be respected. Not because it's legally required, but because they are our peers in this industry and worthy of respect.
And it's all well and good to make those decisions, but if code-generating AI is anything like text-generating AI, they're going to want more. Probably not a great plan to piss off all the people who produce the thing you need, for free, before you have enough of it. Assuming there is such a thing as enough of it, of course.
These “developers” are very immature and naive.
SO is a business
It's always a good idea to read the copyright/usage terms of sites you might be contributing to.
Stack Overflow has been very transparently licensing their content under creative commons since the beginning; that's a big part of why I occasionally contribute to it.
If they allow users to delete or edit their own content, then what's the issue? That just sounds like a technical problem StackExchange didn't think of.
On the other hand, why is this behaviour deemed unacceptable, and StackExchange double dipping on user generated content is not?
I don't agree with how StackOverflow's current management has been running the site, but for anyone who has used the site, the outcome of this behavior should have been obvious. I'm sure Wikipedia would react the same if you deface all the pages you contributed to prevent OpenAI from consuming words you wrote.
At least I'm engaging on StackExchange with the understanding that conversations (questions, answers, comments) there go on the public record for the benefit of future curious humans. I see that as a way of giving back to a community that's provided me with invaluable knowledge.
If I'd known somebody might delete their half of a fruitful conversation as a political statement months or years down the road, I might never have spent the time and effort of participating.
Thinking of it in terms of co-authorship, one of several co-authors also doesn't get to retroactively destroy the combined work by ripping out their individual contributions. If you collaborate, make sure you know and are aligned with the terms of the resulting work.
Great, so how did these companies expect to stab so many people in the back and then continue as if it didnt happen again?
If Stackoverflow is trying to make exclusive deals with Openai, that is against the collaborative spirit of the platform, and I will stop contributing. After all, Openai is charging people for service. If Openai are the only ones given access, Stackoverflow becomes a gatekeeper, peddling my contributions. It'll beget a fork.
I don't think so. Content contributed to SO gets licensed under CC-BY-SA, no? How do these AI models respect the "BY" portion of that license? How do we enforce the "SA" portion of the license on people who use the output of the model?
The answer is they don't, but the big companies decided violating copyright is OK so long as they do it all at once, because there's a lot of money to be made if you ignore copyright.
Personally I'd be OK with a compromise where any entity that creates or uses an AI trained on copyrighted data forfeits copyright on all of their own works (not just those created by AI, all of their works).
CC ShareAlike requires attribution + derivative works to be shared under the same license. ClosedAI is absolutely in the wrong here.
AI is unpopular and makes this an unpopular move, enough so to cause this drama. It's really not any more complicated than that.
Now, as a matter of academic courtesy it's nice to cite the SE post if it's feasible, but it's impossible for me to remember where I learned everything I've learned over the years. Likewise, an LLM shouldn't be forced to cite where it learned facts, which might be technically hard/impossible, but if there is a clear best source then as a matter of politeness it should. And LLMs mostly do this when they can!
Scientific papers have included references for as long as I can remember. It's not about courtesy, but precision and acknowledgement.
So, are we letting LLMs run free and do things that would be frowned upon if they were done by humans?
That'd be like saying my computer learned the film Oppenheimer when bittorrent copied it to my computer's storage system, transcoding it to a lower bitrate, and claim I'm free of copyright law because my computer just learned some facts.
All of OpenAI's products are databases with highly lossy compression applied. Using language that implies a mind is pure scam.
It's less clear that the license gives AI companies the right to train LLMs. They tend to take the view that any media can be used for training regardless of licensing because it's fair use. Many authors disagree and believe they must explicitly license their content for LLM training. AFAIK none of these legal disputes have been settled yet.
In my view you shouldn't really expect to have much control over who uses the content you post to Stack Overflow. Maybe your answer will be used by a CS101 student. But it could also be used by a defense contractor, gambling site, North Korean hacker, or darknet drug marketplace. Yes, the license requires attribution, but if ChatGPT ignores that requirement it's behaving like 90% of programmers.
Like, it's one thing if a fact-collecting business like the NYTimes doesn't want its stories to train an MML. I think under current law they don't have much of a case, because facts aren't copyrightable, but there's a reasonable argument that the law should be updated somehow in light of technological change.
But the work produced by all StackExchange users is explicitly released under a CC BY-SA license. The whole point is to collect and publish facts/ideas/understanding for anyone to see and use for any purpose, including running a business. Yes, the "SA" (share alike) part means if you want to use and modify the words then you need to release them under license that is at least as permissive, but LLMs aren't using the words; they are clearly digesting the facts and expressing them in their own words. And, unlike the NYTimes, there is no issue of "couldn't new tech undermine society's current method of economically incentivizing fact-collection?". The StackExchange users are not being paid, and the fact that the license is not NC (non-commercial) explicitly means that using their hard work to make money is allowed (and encouraged!).
The CC BY-SA text <https://creativecommons.org/licenses/by-sa/4.0/legalcode.en> says nothing about the words. The words "word" and "words" do not even appear in the legal text. What it says, however, is:
> In addition to the conditions in Section 3(a) , if You Share Adapted Material You produce, the following conditions also apply.
> 1. The Adapter’s License You apply must be a Creative Commons license with the same License Elements, this version or later, or a BY-SA Compatible License.
> 2. You must include the text of, or the URI or hyperlink to, the Adapter's License You apply. You may satisfy this condition in any reasonable manner based on the medium, means, and context in which You Share Adapted Material.
where "Adapted Material" is defined as:
> material subject to Copyright and Similar Rights that is derived from or based upon the Licensed Material and in which the Licensed Material is translated, altered, arranged, transformed, or otherwise modified in a manner requiring permission under the Copyright and Similar Rights held by the Licensor. For purposes of this Public License, where the Licensed Material is a musical work, performance, or sound recording, Adapted Material is always produced where the Licensed Material is synched in timed relation with a moving image.
So if you want to argue that LLMs don't have to follow the SA clause, then you have to argue either that:
1. LLMs aren't an arrangement or transformation of their input, or
2. the license doesn't apply at all, eg. because it's fair use in whatever jurisdiction they fall under.
OpenAI is using argument 2 <https://news.ycombinator.com/item?id=37780199>.
In other words, it has nothing to do with your or anyone else's interpretation of what the CC license claims to protect ("an arrangement or transformation"), and only relies on what rights they are legally allowed to assert.
Injecting an LLM middleman between your post and the end user changes this dynamic quite a bit - without the human component, the feeling is that you're just doing unpaid labor for a profit-oriented company (OpenAI).
It's not about legality, it's that sites like SO and Reddit have turned a corner and are no longer sites or businesses that people want to support. The contract of user generated content has been disrupted, we'd have happily given them content to profit from forever if they didn't turn shitty. And now people are lashing out even if it's futile because the content has already been sold. These sites failed a generational marshmallow test, taking short term profit at the expense of souring people on the idea of providing free content going forward.
Without the community posting and answering questions, ChatGPT won't work anymore. They (SO) have to block any scrapping, wait until GPU prices go down, or use open-source LLMs with their data and figure out how to monetize that service (maybe giving some percentage to devs, etc...) or even check if that approach (add a chatbot service) makes sense.
They just caved to the FOMO mentality without considering that tech is always evolving and they need engineers, dev, etc... to keep finding out bugs, writing about their experiences on how they solved those errors, etc...
We'll see if a different contender figures this out and comes out with that solution if SO doesn't change its course.
The idea of Pillaging the Commons is interesting. I wonder when the mainstream opinion started shifting. It wasn't quite sudden but it seems to me that even ten years ago the dominant Internet visible position was that much of copyright was bogus: information wants to be free / if a pirate copies your stuff, you still have it
But now there's a stronger sense of "this information is ours". Perhaps that subculture moved somewhere and this one came here or perhaps I moved from where the former was to where the latter is.
I find it an interesting sociological phenomenon.
If anyone else thinks it's a bad move, the most efficient way to boycott it is by adding senseless questions and answers and upvoting them. Us people is how they got big, us people is how they go down.
I know that SO content is licensed under creative commons; is there an additional dual license that allows them to commercially use it under different (e.g. non-attribution) terms?