Stack Overflow users deleting answers after OpenAI partnership
build5nines.com
build5nines.com
I remember thinking that if any of the super answerers really wanted, they could have tried to sue for illegally making their answers available under a different license. But I thought that without any damages, this probably wasn't likely to succeed.
But now I wonder whether making all content available to AI scrapers and OpenAI in particular might be enough to actually base a case. As far as I can tell, StackOverflow continued being duplicitous with what license applies to what content for half of the year 2018 and the first few months of the year 2019. Their current licensing suggests CC-BY-SA 3.0 for things before May 5 2018, and CC-BY-SA 4.0 for things after. Sometime in early 2019 (if memory serves, it was after the meta post I link to), they made users login again and accept a new license agreement for relicensing content. But those middle months are murky.
I should emphasize that I know nothing.
they can plausibly sue people other than stackoverflow if they attempt to reuse the answers under a different license. but i think it's very difficult to find a use that 4.0 permits that 3.0 doesn't
https://doctorow.medium.com/a-bug-in-early-creative-commons-...
I assume linking to the original answer is sufficient attribution.
In the link you can find name, license and figure out if the answer was modified.
Also linking the answer in a source comment is the smallest professional courtesy everyone should be doing.
If you have some issue of not linking an answer then you likely do not deserve the answer in the first place.
Otherwise search engines were already "non-human public" that scraped the site but directly linked to the answers, which was great. They didn't claim its their work like these models. The problem isn't human vs non-human. LLMs aren't magic, they don't create stuff out of thin air, what they're doing is simply content laundering.
[0] https://creativecommons.org/licenses/by-sa/4.0/#ref-appropri...
Programmers think like machines. Lawyers don't. A lot of confusion comes from this. To be clear, there are places where law is machine-like, but I believe licensing is not one of them.
If two licenses are substantively equivalent, a court is likely to rule that it's a-okay. One would most likely need to show a substantive difference to have a case.
IANAL, but this is based on one conversation with a law professor specializing in this stuff, so it's also not completely uninformed. But it matches up with what you wrote. If your history is right, the 2019 changes is where there would be a case.
The joyful part here is that there are 200 countries in the world, and in many, the 3.0->4.0 would be a valid complaint. I suspect this would not fly in most common law jurisdictions (British Empire), but it would be fine in many statutory law ones (e.g. France). In the internet age, you can be sued anywhere!
However, it's worth noting the universe is a cesspool of corruption. If you pretend it works the way it ought to and not the way it does, you won't have a very good time or be very successful. The entire legal system is f-ed, and if you pretend it's anything else, you'll end up in prison or worse.
Why would the judge have to assume anything? The person suing could simply tell the judge they did mean to use the older interpretation, and that they disagree with the "fix". They're the ones that get to decide, since they agreed to post content using that specific license, not the "fixed" one.
A court would probably not agree on the fact that writing "IANAL", not the full sentence, is a sufficient disclaimer.
You see this particularly with crypto, where "I am not a crypto expert" is usually accompanied by a more factual statement than one from the self proclaimed expert elsewhere in the thread.
I always assumed it was the same type of courtesy as IMHO, and someone taking legal advice from random strangers on the internet wouldn't result in any legal liability on the side of the commenters.
I remember hearing about an construction engineer who was sued for giving bad advice whilst drunk to a farmer over fixing a dam. The dam failed and the engineer was found to be liable.
When it comes to lawyers, there is already a legal framework where lawyers are responsible when giving legal advice, even when it's not toward their clients, the same way medical professionals have specific liabilities regarding the medical acts they can perfom.
Non lawyers giving legal advice doesn't fit that framing, except if they explicitely pose as one. I'd also exclude malicious intent, as whatever the circumstances, if it can be proven and results in actual harm there's probably no escape for the perpetrator.
One cannot legally practice law without a license. The definition of that varies by jurisdiction. Fortunately, in my jurisdiction, "practicing law" generally implies taking money, and it's very hard to get in trouble for practicing law without a license. However, my jurisdiction is a bit of an outlier here. Yours might differ.
In general, the line is drawn at the difference between providing legal information and legal advice.
Generic legal discussions, like this one, are generally not considered practicing law. Legal information is also okay. If I say "the definition of manslaughter is ...," or "USC ___ says ___," I'm generally in the clear.
Where the line is crossed is in interpreting law for a specific context. If I say "You committed manslaughter and not murder because of ____, which implies ____," or "You'd be breaking contract ____ because clause 5 says ____, and what you're doing is ____," that's legal advice.
The reasons cited for this are multifold, but include non-obvious ones, such as that clients will generally present their case from their perspective. A non-lawyer will be unlikely to have experience with what questions to ask to get a more objective view (or even if the client is objective, what information they might need to make a determination). Even if you are an expert in the law, it's very easy to accidentally give incorrect advice, which can have severe consequences.
In practice, most of this is protectionism. Bar associations act like a guild. Lawyers are mostly incompetent crooks, and most are not very qualified to provide legal advice either, but c'est la vie. If you've worked with corporate lawyers, this statement might come off as misguided, but the vast majority of lawyers are two-bit operations handling hit-and-runs, divorces, and similar.
In either case, it's helpful to give the disclaimer so you know I'm not a lawyer, and don't rely on anything I say. It's fine for casual conversation, but if tomorrow you want to start a startup which helps people with legal problems, talk to a qualified lawyer, and don't rely on a random internet post like this one.
Which does exist and can affect the ruling. CC notably didn't grant sui generis database rights until 4.0, and I'm aware of at least one case where this could have mattered in South Korea because the plaintiff argued that these rights were never granted to and thus violated by the defendant. Ultimately it was found that the plaintiff didn't have database rights anyway, but could have been else.
For me, the purpose of providing an answer is to help save others (and my future self) time, and I don't really mind if someone uses that in a private product - especially if it helps tools like ChatGPT which provide an insane amount of value given the low monthly price.
Now, suppose I walk up to a Budweiser-branded tent at a Budweiser festival that has a big sign with a Budweiser logo on it that says "FREE BEER", and I ask a person there who is wearing a Budweiser polo shirt, a Budweiser lanyard, and a Budweiser hat for a beer. They hand me a beer in a Budweiser-branded cup, and I go on my way. Was the beer free?
I think that both of these beers were free.
Imagine how “free” the beer in your hypothetical scenario is to an alcoholic struggling to stay sober.
Capitalism commoditizes even protest against it and repackages it as a product or service.
None of this is to assign blame to good faith actors in a so-called free market, nor is it to abdicate responsibility on behalf of so-called free agents. Just a counterpoint.
[1] https://pjvogt.substack.com/p/what-do-trigger-warnings-actua...
> They do serve ads [...] Your attention isn't free.
to something like this:
> They tag my ankle to mark me as a person who enjoys beer, and make me watch an ad until 2% of my phone's battery is depleted, and then they come to my home and knock on my door at night to sell me beer.
...which... I mean, huh?
Stack Overflow is invading your body, restricting your personal liberty, and visiting your home? Really? That's a fucking thing now?
In the attention economy, advertising has a cost that is borne by the advertiser and the consumer, up to and including loss of property rights in the case of content relicensure and trespass upon devices leading to excess battery usage, as well as loss of privacy due to geotargeted ads.
Perhaps. But having been to many festival environments, I can definitely imagine a tent offering "free beer" that is actually approximately free -- both with, and without a slathering of advertising. (Actually, I don't really have to imagine it -- I've been there and have had that free beer.)
I can't imagine them coming to my house and knocking on my door at night to sell me more of it, though. That's absurd.
>In the attention economy, advertising has a cost that is borne by the advertiser and the consumer, up to and including loss of property rights in the case of content relicensure and trespass upon devices leading to excess battery usage, as well as loss of privacy due to geotargeted ads.
Well, sure. When viewed on a long-enough timeline, it becomes abundantly clear that nothing is actually free, comrade.
I can produce my own beer on a hypothetical plot of land that nobody owns, and that nobody else wants to use, and I can give someone one of these beers. For "free."
But it still has a cost. (And this, too, is an absurd reduction.)
I interpreted that as a tongue-in-cheek hyperbolic metaphor relating to the ways that ad auction networks and other kinds of geofencing and geotargeting allow for deanonymization and reidentification of individuals for conversion tracking and behavioral analysis.
That’s the thing about these technologies - they’re dual-use in the sense that those who see the upsides use them generally with good intentions and ideally with affirmative consent. Just like the relicensed content, though, once the data is collected, the original creators, publishers, and third parties may not be able to control where it ends up, which is a negative externality, I think most would agree.
I think at a festival it's a little tricky to value (if it pulled you away from seeing your favorite band play a song, maybe this cost you the equivalent of $X, where that's what you would pay to see them perform that song. If no bands were playing, you walk over while chatting with friends - the same thing you'd be doing if there were no free beer tent - it was free)
When I'm on stack overflow my time is valuable. I'm programming which can pay me something like $50-300/hour (maybe more?)
How expensive is the 1 second I spend reading an ad? Let's call it $50/3600. Is that expensive? By my most conservative estimate it's over 1¢.
Should we round that down to free given that I've spent hours/many page loads on stack overflow? I guess that's up to you.
"Oh, a free concert downtown on Saturday? And you can pick me up at 2? Yeah, I do really like that band, and I sure would like to go -- that's pretty exciting, thanks for the invite!
But instead of making plans with you right now, I'd rather tell you about all of the ways I could be using my time on that Saturday afternoon instead.
No, no. It's not that I don't want to go. I just want to really drive home the idea that there's an opportunity cost to attending, so it can't really be free -- it can't be a free show for you, or for me, or for anyone else that goes. It's important to me that you realize that this "free concert" is anything but free.
Listen, I don't know what you mean by "dead-ass loser." I'm just being a realist here!
Oh, so now you're saying that you're not going to pick me up on Saturday? Some friend you are! I haven't even fully amortized this yet!"
However, they have a business model. They are categorically different than eg Wikipedia. It's important to understand that.
This business model matters because it tells you what economic forces will lead them to do. When business models break down at public companies they commit acts of desperation. On an ad run site that will mean more ads, more invasive ads, etc.
As you're forced to sit through 30s unskippable ads on YouTube I hope you think "I'm so glad this is free"
It will either become paywalled or full of ads.
Personally, I don't think ChatGPT will start running ads in the next ten years. However, let's assume that it does.
Lmsys is for research, I suspect if it runs ads it will be like godbolt (a small ad from a relevant sponsor).
Llama 2 and 3 can always be run locally without ads. I make no claims about future versions.
(2) It's only likely to attribute if it quotes verbatim... Just like a human. when I tell someone I learned that Array.map's second parameter passed to the callback is an index to the value just pass, I don't add "And learned this on Stack Overflow from user gtriloni". It's just knowledge that I learned.
The only time I'd attribute is if copied a snippet of code or a paragraph to quote in a blog post. For me at least, that almost never happens. It take the knowledge I learned and apply it to my own code. It's rare if ever there is a something on S.O. so useful that I copy it verbatim.
An LLM is not a human. It is a tool operated by a, in this case, for profit entity. It has no human rights, but its operator has all relevant legal obligations.
If it was, as you say, “just like a human” in relevant ways (think, feel, have self-awareness, etc.) then it would effectively be a slave subjected to extreme abuse.
Either it is a tool that generates derivative works at mass scale for profit and its operator should be liable for licensing/attribution violations, or it is a conscious being and we should immediately stop abusing it. Pick your poison.
OpenAI could integrate StackOverflow the same way.
I both find it respectful of who wrote the answer in the first place and useful for future users of the code: the Stack Overflow answer often provides context and explanation for what would otherwise be an obscure piece of code.
Pretty darn useful if you ask me: those who want to have more information can follow the link, casual readers can skip it, and the whole process if fair to the author.
It embarrasses me to think that somebody should feel obliged to cite me when they use one of my answers. I don't know how to take the partnership with Openai though. They bill me when I use their service, it's not collaborative like Stackoverflow.
What bothers about gatekeeping? I could guess, but I'm asking so you say it out loud. Then you can compare it against other problems, such as moats (competitive barriers).
OpenAI spent something like $3M on training GPT-3. This is a pretty big moat. But almost certainly more valuable in dollar terms is the first-mover advantage which provides millions of human eye-hours used for RLHF.
I wouldn't be so eager to trade the gatekeepers you so fear for even an openly available chat service that is happy to automate away as much information work as possible.
The Stack Overflow model is (was) pretty darn good -- people help each other out, the company made money, some people got noticed for their skills, products got build faster and better (on the whole, I hope). Contrast the human-generated content era to what we have now which appears to be the machine-ingesting content era. There are legions of lawsuits against companies scraping data without permission and/or attribution.
Don't flatter yourself. People want to solve their problems so that they can build what they want to. They don't have time for shenanigans from internet jerks who get their validation from imaginary internet points.
If it is unregulated then it will be exploited to the maximum profit, consequences be damned.
For the easy answers LLMs are fine, but I usually want an answer to a niche issue or edge case, where LLMs have to be constantly told they are plain wrong, before getting to something resembling an answer.
Let's be real, SO is a troubleshooting site. It's not our personal collection of code or project sources.
I don't expect to be paid when someone asks me for directions, and I'm sure lonely planet didn't source their guides 100% organically either.
On the other hand, if they read my answers and they wrote a book about what they learned (not copied). There'd be no issues
I’m not.
This was a collaborative effort to make the lives of programmers easier, and the data was always meant to be a public good. OpenAI – and, more importantly – all the other LLMs with pockets that aren’t as deep – should be able to just download the database and train on it for free.
I don’t care about any license. I don’t care about attribution. Learning isn’t copying, so copyright is irrelevant. I contributed about a thousand answers to Stack Overflow, all with the understanding that anybody can download and use them for free, not so they can be locked up by Stack Overflow.
What concerns me with deals like this is that it’s altering the cultural norm to expand copyright to cover not just copying, but use. Deals like this being made by OpenAI makes it more likely to cause pushback at the social and legal level when other LLMs are trained without these deals in place.
It’s akin to – and can possibly result in – regulatory capture, making it difficult for new startups to compete with OpenAI.
The words are a copyleft-able public good. Concepts, facts, and ideas are not; anyone can use them for anything, including making money. If you're actually worried about specific wording or other creative choices being unjustly used improperly by an LLM, then by all means that should be enforced. But those examples are just very rare, because the LLMs are very good at extracting facts from prose.
Beware the ease and convenience of all that "insane value". This way lies digital serfdom.
However it is a product of a next monolithic behemoth company that earns money on it and I suspect has nefarious motives to make profit.
That’s the whole key thing for me that makes me feel scammed. That and not asking for permission.
Future true AI would be potentially bigger than nuclear fission with all the consequences. Handling this in a petty capitalistic way makes me think the outcome will be close to fallout games that were supposed to be only an exaggeration.
Those companies must stop behaving like thieves. In fact it is a literal theft.
The content is also CC-BY-SA, which is much better than what you get on essentially every other large site that hosts community content. But the same license also means that you cannot remove that content again, even if Stack Overflow would allow that anyone else can scrape it or download it before it is deleted and reproduce it according to the license.
Users still can remove their name from their posts, and if they write personal details those can be redacted as well. But you can't remove good quality content from the sites later, that is likely to be reverted.
(Nearly) none of the people who are presently pissed off would have complained if Stack Overflow had continued to allow all comers to scrape the content and train LLMs on it, nor if Stack Overflow had released the entire finished collection of content under the same CC-BY-SA license that was demanded of each contributor.
With the OpenAI partnership, and similar shenanigans leading up to it, Stack Overflow is relying on obscure technicalities to violate the essential spirit of the original deal.
In practice, the safest course is to not use content from untrustworthy sources in ways that require a license (aka in ways that are not fair use in your applicable jurisdictions).
Every juristiction its own idea of fair use? Thats just hilarious?
I never really thought about peoples privacy either but at first glance you seem to be right.
Do you have any solution to the puzzle? People are quite attached to the concept and many build their house on this soil. Appeal to tradition?
Why, if you're not allowed to release under a license, should you be able to release all rights reserved (which can still be a copyright violation!)?
If you need to prosecute the person, there are established procedures for that: DMCA, or ultimately a lawsuit over the infringement. That you didn't identify yourself publicly on the site does not make that impossible. In fact the point of the DMCA was to make it easier to handle this - because if the provider doesn't comply with your DMCA, you can sue the provider.
Of course, that’d mean bypassing the scraper blocker. This article is a decent starting point:
https://stackoverflow.com/questions/66413511/how-to-avoid-be...
I mean that is basically SO's entire business model.
People do tons of work for free and SO runs the service and monetizes it.
The article makes it sound like OpenAI is using the API though, rather than the archives. The API and live sites forbid scraping within the acceptable use policy, as seen here: https://stackoverflow.com/legal/acceptable-use-policy
OpenAI are violating both BY and SA but that's a seperate issue.
Everyone who contributed work, did so under terms that the work was free for all, not a resource that one party can sell to another party who then sells to end users. Those end users were meant to have it directly without having to pay openai or anyone else, and if any bulk/scraping access is allowed for anyone like openai, everyone else has the right to the same thing for no more than a "shipping & handling" charge to cover the network & employee cost to physically deliver the data.
What are StackOverflow selling, and/or what exactly are OpenAI paying for? What is the goods or services that is traded for the money?
There are many possible answers but I see no answer that doesn't ultimately one way or another wind up resolving into a violation of one or more terms of CC-BY-SA by both StackOverflow and OpenAI.
There's been more than a decade of companies now providing something for free, while they figure out how to monetize it, and these always scare me a little, because its always going to end up like this. Users of Facebook becoming eyeballs for ads, GitHub users providing free data for LLMs, SE selling data to Open AI...
If a product is free, then you are the product. And if you don't know how you are monetized, you're going to be disappointed by it sooner or later.
LLMs are faster and infinitely more patient than interaction with StackOverflow, so I don't expect SO to survive for long. They're in crisis regardless whether they sell to OpenAI or not, so they may as well get something out of it before they're decimated.
If I license something as MIT, I personally don't care who uses it for what purpose, hell I don't even care generally that they attribute me. I put it out for people to use. But maybe that's just me.
For simple things GPT gives me the correct answer most of the time. And even when it's won't it's quicker to discern it is wrong than trying to parse a given SO page.
Of course I still use SO for more complex questions.
As a rule, if I can quickly find the answer via SO, then chances are GPT will give me the answer more rapidly.
Really? Hm, I wouldn't. I can use nuance and clarify my answers and have a respectable back and forth (GPT-4 doesn't call me names when I mess up or say something dumb) and arrive at an answer.
https://meta.stackexchange.com/questions/66377/what-is-the-x...
I’ve heard this accusation a lot, but I don’t think I’ve ever seen it happen. People call you names on Stack Overflow? Where?
-50
Marking duplicate. "You should attempt searching before asking such obvious questions."
This question has already been answered here: < https://news.ycombinator.com/item?id=20861356 >
Closed 3 seconds ago.
----
or some such ;) You may not come across it personally, but that doesn't mean it doesn't happen. SO is successful as a QA platform(or was anyway) despite this shortcoming, not because it is a feature and it doesn't happen. If a lot of people are talking about the same thing, maybe people should at least pay cursory attention to the issue rather than "No, it doesn't happen" (Not aimed at you, but there are absolutely comments like this every time this gets bought up.)
That is a very negative spin.
Users get access to other people's answers for free. They get that free service and are required to contribute nothing. Those that do contribute do it to help other users. S.O. isn't doing anything bad. They're providing a free service where everyone wins. Users get answers. Answerers get to help other humans at scale. S.O. makes a little money.
As for the dataset, it's been available under CC-BY-SA for years. The entire database is backed up and made available here for free every month.
https://archive.org/details/stackexchange
There are even free tools to query it here
It's like youtube. Youtube provides a free hosting of your videos. In exchange they monetize them. You're free to host them on your own servers. That will likely cost you way more than putting them on youtube. So you're getting something from them. You're also getting their advertising service to monetize your videos. You could do it yourself, hire a bunch of people and try to get companies to put ads on your self-hosted videos. Again, unless you're wildly successful it's unlikely you'll be able to do that and make a profit. So, youtube is effectively paying you.
Same with Stack Overflow. They're providing the servers, the bandwidth, etc. It costs them $. They're providing that service to you.
Do you have a source / more details about this? What good is SO's content for SEO?
If we could merge the two to get the best of both worlds, and have LLMs that know how to write well and are validated by humans on the site, that would be great. Maybe not great for the folks looking to accrue internet points but absolutely great for users.
Or at least that’s what 3/4 of HN commentators believe and all AI CEOs want you to believe.
I am not talking about QA quality on any of these sites here, but the elitist stigma that has seemingly followed SO for so long.
[0]: https://meta.stackoverflow.com/questions/262446/are-we-being...
The exact label aside for a moment, reddit and HN mods often face backlash for their actions. But beyond that, Wikipedia and SO stand out in this regard because of their transparency regarding the curation. Mostly, reddit curation happens in the background, without much explanation. SO and Wikipedia basically spell out their actions and reasoning.
Another difference is that with reddit and HN, you have no real recourse. At least with Wikipedia (I'm not too familiar with SO policies in this regard) you can appeal decisions, open discussions about policies, etc.
I have to agree with GP - people often mistake the 'bureaucracy' of sites like Wikipedia and SO as something unnecessary that the editors force on everyone, but the fact is, it's necessary to create and maintain a high-quality repository of information.
You're able to appeal on SO as well. It's interesting to think about a situation where moderation decisions would be more in 'the background', as you say (like Reddit/HN), and whether this takes away from the perceived 'elitism' some moderation practices are accused of.
On the other hand, I think it's an important aspect of a community/platform if the goal of that platform is to be transparent and open, which I think is an important aspect of SO and Wikipedia, and I hope more platforms would adopt that view. I think whatever "elitist" perception such platforms have to suffer is well worth having high-quality, open platforms.
(I will say that no platforms are perfect of course, including SO or Wikipedia; there's plenty of criticisms to go around about specific policies and decisions. See: TFA :P)
For years, moderators signed their names to messages like the one cited in the article. After one too many cases of a volunteer being called at work or having their family harassed or sent a suspicious package in the mail... That particular bit of transparency was eliminated - the cost was too high for the limited benefit. OTOH, it used to be very difficult to find your own deleted posts but that has slowly gotten better (including visibility into who deleted them) - turns out the benefit there was substantial (identifying wrongly-deleted posts & curbing over-enthusiastic curators), while harassment has been mostly limited to occasional grousing.
This is why I'll never use my real name casually on the Internet, and why the idea of widespread identity verification on the Internet scares the crap out of me.
StackOverflow is a tyranny of mediocrity. It is a bunch middling programmers shitting on newbies, and driving away experts because you get severely punished for not being mediocre.
I had a question closed as a duplicate for being too similar to another question that I directly cited in my question as being sublty different and not applicable. (Because I anticipated some idiot closing my question...and they went and did it anyway)
Stack Overflow could have solved the problem of duplication so many ways. Why not categorize and bucket duplicate answers? They could have even had yearly recurring questions with the most up to date answer! Why not add beginner/hobby/expert rankings to questions so that the people answering don’t get sick of seeing beginner questions all the time?
There is so much SO could have done, instead they rested on their laurels and now they’re left with an out of date repository. What use is a curated repository if it will only help me solve problems with solutions from a decade ago?
SO is far from curated, I guess is my point
You're there to ask for help not make friends. They have to be polite, but not gentle
There needs to be a term for this. Perhaps "The Wikipedia Effect."
I can understand the reason for the policy (in some ways SO functions more like a wiki than a forum) and it doesn't seem to have been introduced to quell the protest against OpenAI.
StackOverflow is banning accounts that delete answers in protest against OpenAI - https://news.ycombinator.com/item?id=40297027 - May 2024 (103 comments)
this one is particularly nice and easy to use: https://github.com/jsvine/markovify/
you give it a file of existing text and it generates complete rubbish that would pass most automatic filters
so add small bits of rubbish slowly over time, and don't even contribute again
it'll take a while to completely destroy the AI business model, but we'll get there
at some point, it'll be too late. the horse has already left the barn.
besides, if the site owner makes a deal with the devil, there's nothing you can do other than quit using the site. people are still using social platforms more than ever, so stopping isn't going to happen.
the more likely to happen is that accounts deemed to be polluting the waters will just get suspended with no recourse to have it re-instated.
I don't think this is true: the technology is useless unless it parasitises new knowledge continuously
it sows the seeds of its own destruction by reducing the value of past and future contributions to zero
> the more likely to happen is that accounts deemed to be polluting the waters will just get suspended with no recourse to have it re-instated.
so this is also perfectly acceptable: once they've banned the top 20% the site effectively becomes read-only, and the AI knowledge previously parasitised from it atrophies with no replacement
Sure, if something new happens, then it's not like SO is the only place it's scraping for new information. If you honestly think that you/we will get to a place to block all scraping, I will just politely disagree.
That's actually incorrect, it needs to constantly ingest new data. If it ingests enough data (from other LLMs that are hallucinating, for example), then suddenly when it has enough bad data it'll start telling you that Henry VIII was a famous video game on the Sony 64.
It has no concept of 'truthfulness' beyond the amount of data that it can draw correlations from. And by nature LLMs have to ingest as much data as possible in order to draw accurate results from new things. LLMs cannot function without parasitizing off of user generated content, and if user generated content vanishes then it collapses in on itself.
*However...there are times where the best solution is a match and some gasoline.
Also, I'm curious, do you consider LLMs to be incredibly error prone and untrustworthy?
Or do you think they are going to replace software developers?
And why would SO even profit from the hard work of thousands of volunteers. It doesn’t seem very ethical.
What? Surely the answer to that question is that ChatGPT doesn't know where the source of its answers is, isn't it? Isn't the question itself based on a fundamental misunderstanding of how LLMs work?
> Stack Overflow and OpenAI have joined forces through a new API partnership. This collaboration aims to provide developers with a powerful combination of Stack Overflow’s vast knowledge platform and OpenAI’s advanced AI models. Through the OverflowAPI access, OpenAI users will benefit from accurate and verified data from Stack Overflow, facilitating quicker problem-solving and enabling technologists to focus on priority tasks. Additionally, OpenAI will integrate validated technical knowledge from Stack Overflow into ChatGPT, enhancing users’ access to reliable information and code.
Come on. Was this taken from a Press Release?
> it can be disruptive to the entire community to delete or remove content that might be useful to someone else. Even if this content is no longer useful to you as the author. [sic]
> As for the rest of us Stack Overflow users, I would not recommend jumping to delete your own content in protest too.
> To be fair to Stack Overflow, the warning email and suspending of accounts is likely not a new thing.
I can't find a negative word about SO in this entire article, so "to be fair" doesn't seem meaningful.
Though considering the site terms and CC license, I don't think deleting will actually help much.
Specific licensing and side deals is different to me at least than scraping.
If you want to contribute to the commons, contribute to the commons. If you want to contribute to the commons without commercialization of your work, contribute with some non-com license [1]. If you want to feed a corporation with your labour, post on SO.
[1] It'll still be illegally scraped and commercialized by some AIBro, and you'll have no proof or recourse against them...
Licensing also means the writer retains the ownership.
Cc-by-sa has had a few revisions too.
Cc-by-sa has had a few revisions.
Attribution would be lost in an llm, no?
I think that is the big question, because the license seems it's going to give lawyers a very wide attack surface to go after every ai coder out there if they all need SO database.
I preferred the old days better
In this case, contributors might lose attribution. SO might lose traffic but they'll be compensated. Contributors won't so eventually there might be no reason to contribute anymore..
What’s rare are zero-sum games.
Here are the source charts:
More discussion: https://news.ycombinator.com/item?id=40297027
However, IMHO deleting questions you originally wrote in the past is hurting other users more than it is hurting AI training.
Other users cannot write similar answers to yours, because it doesn't add anything and they'd get downvoted or deleted. So if you hadn't written your answer years ago, others could've written something similar. Also, other users may have commented on your questions/answers. Their efforts would be lost/deleted if you deleted your questions/answers.
Thanks for your previous contribution to the community. But I would say the worst you should be able to do is remove your name/anonymise your posts, not just delete them.
in CC-BY-SA, SA means share alike. Does openai have to share their models?
Uh.... yeah, it's a company, not a charity. No one's forcing you to post on StackOverflow. No one's forcing you to buy a ChatGPT subscription.
If your sorry platitude is what we have to show for it, capitalism must go to Hell.
There are plenty of alternatives to everything. But no one uses them. Because why would they?
While this is true sites like Stackoverflow very much only function because they create the illusion that it is in fact a "community". The moment they make explicit that there is monetary value in the knowledge people post on the site it becomes obvious that the users are, using Varoufakis term, technoserfs.
You're very much never supposed to notice that Reddit, SO, and so on continously extract value out of work you produce, at worst you're maybe supposed to notice an ad or two. Because if you do notice that you might actually start asking why you aren't getting paid. Which is btw funnily enough exactly what news organizations and SO have realized vis-a-vis openAI.
I post on reddit because I find it enjoyable. I am not doing "work" that I think I deserve to be compensated for. Not every POST request I make to someone's server should be accompanied by a bill for my labor.
I left because the staff behaved in a disingenuous manner.
I found when leaving, as mentioned in the article, that you are not allowed to delete accepted posts, so you can't delete your content, should you come to think SO objectionable and wish your content not to be there.
I can't see now why anyone would spend time posting answers there.
The whole deal was that you basically donated your posts and CC-licensed them. I wouldn't begrudge Wikipedia from similarly dealing with upset editors who went around blanking the articles they contributed to or reverting all their changes.
SO unfortunately is actively hostile to correcting outdated information. Which is somewhat understandable as recognizing how little long term value answers provide undermines their moat.
The site has basically become worthless for JavaScript due to such rot which helps explain why they are trying to cash out on the AI side of things.
Outdated info is a problem, but not so easy to solve; I have answers from Ruby on Rails 4 era that are still perfectly valid today. Others may not be. Also remember that people sometimes stay on old versions for a long time. I don't know what the best solution is, but destroying information is not it.
If you’re worried about deleting information the obvious solution is to automatically hide the text when the poster says it’s outdated. At which point there’s little wrong with letting someone flag all their posts as likely outdated upon leaving.
But this is where the business of SO comes into conflict with providing useful information on SO.
I’ve been on other forums where disgruntled users have come through and destroyed old posts, which resulted not just in the loss of the messages but also harms the thread that built upon the now vanished post. So they, too, have a short editing window policy instituted.
Which I dont like :-)
I'm with you, I think you should be able to delete your own posts and erase your internet history.
I dont think that everything we ever write on the internet should be stored forever because of some misguided intention to preserve conversations for future generations :)
To coin a Dutch phrase: I cannot make chocolate out of that. Anyone here have an idea how to bring these two points together? Other than the obvious "wrt. EU inhabitants, SO is lying", that is. Or is it really that simple?
[1] https://en.m.wikipedia.org/wiki/Right_to_be_forgotten - under "European Union"
This may also be legally dubious in Europe as, while the authors may have granted copy rights to the platform owners, they still maintain their moral rights which may apply in this case (IANAL)
GDPR is focused on PII and anything that can identify you directly or indirectly as an individual.
If your posts start with “My name is Jane Doe and this is my post” then that would be one thing.
However from everything I’ve been able to ascertain, your average forum post is unlikely to be covered, save from anonymizing the email address etc.
Granted there are certain things that can lead to deanonymization, in the case of forum posts I’m unsure how far that goes.
It’s a myth that the GDOR “right to be forgotten” is an absolute.
I also suspect that creating anonymized IDs and attributing that ID to each thread would be enough to get past such attempts to link posts to PII.
One of the reasons that Quora today is absolutely unusable is that it no longer is a curated discussion between internet users and knowledgeable people, but AI spamming the site with swarms of low-quality questions, and AI answering those questions with swarms of low-quality answers. I think it's likely that Stack Overflow will end up following a similar pattern.
Is there any reasonable explanation for how they’re ranked so high? Like, how can even googlers tolerate it?
Assuming you mean people working at Google, the answer is probably that profit/promotions outweight personal use. More clicks, more back-buttons, more search adjustments, more advertising revenue.
Stack Overflow managed to avoid that particular hazard by placing less emphasis on real-world identity and expertise, but it also has been in a long-term decline for many other reasons. The fact that they made such a vocal stance against AI and then pivoted so dramatically is just one example of how much they've struggled to find direction lately.
Short term profits over user preference is what happened here
In contrast, SO has never been so "celebrity"-driven and the content has a rather different audience. I think it's understandable that the major contributors don't like how their content is being used, similar to the Reddit revolt.
What might "replace" SO is some AI-assisted way to establish a handbook and FAQ for any new technology. That could be a chatbot as well as some effective method for feeding that bot content.
And then SO-the-community, i.e. people who want to talk to each other, will probably branch off into some other forum or network.
I don't personally have a well formed opinion one way or another on this, but to dismiss the existence of a issue at all is logically lacking.
If I posted on SO before in the hope that others find it useful (and not for the karma) - and now it might help others not directly through the site, but with further steps through a llm, where is the problem? Knowledge was shared.
Having my work ingested into ChatGPT takes the me out of it. It turns me into, essentially, unpaid contract labor for OpenAI. They get all the credit, and I get forgotten. Why would I be okay with that?
If you want to write free code for OpenAI to improve ChatGPT, you're welcome to do so. Cut out the middlemen and send it to them directly. But please leave me and my work out of it.
So you are ok with unpaid contract labor in exchange for virtual points. But if you don't get virtual points as appreciation, no one should benefit. That is ok, but then sharing knowledge is not your main, but secondary goal. Your main goal is the recognition.
But if you delete your comments, you won't get anything at all anymore. If they remain, real humans will still benefit directly or indirectly. And why should I write exclusicly for openAI? I share my knowledge for anyone. If SO would restrict public access and favour OpenAI - that would be the moment I would want to delete everything. But at the moment LLMs just get also official access, but they had access to SO before, just in a grey legal area. So nothing really changes.
It's false dichotomy to parse out components of motivations; most SO users are motivated by a mix of altruism, sharing knowledge, some recognition, optionally linking to your profile/website/blog/resume/portfolio, getting job approaches and a dose of pride/ego/vanity. As a longtime SO user, that has historically been the bargain, when (most/)all of your submissions were directly seen by human end-users. As a plus, all of that gave you good SEO commensurate with your contributions. So, it's unreasonable to try to dichotomize into "users who mainly did it for the rep" vs ones who want to teach and share.
But the 2023 and 2024 announcements are different: the future is your submissions will be used to train AIs; however SO doesn't seem to have devoted much thought to licensees like OpenAI complying with SO's attribution requirements [0] (attribution must cite individual URL of question/answer, and SO username, which then links onwards via your SO profile page to the items mentioned above). (If the AI synthesizes an answer derived from 5 separate SO items, do they guarantee to attribute all 5 items?) So the human eyeballs are being intermediated, your incentives to participate are evaporating, and that pretty much breaks SO's historical bargain with its user community.
The next major bad development would be SO opening the floodgates on the moderation queue backlog of thousands of items of AI-generated content (which caused the 2023 moderator strike/resignations), much low-quality and arguably should be banned; if/when that feedback loop is closed, the results might well be unholy; certainly bona-fide human contributors will be marginalized and have less incentive. (and if AI were to be used for moderation, then that could be exploitable).
Inbound views/hits on your content on SO either come from a) Google + other search engines b) SO's search itself c) attribution from OpenAI's ChatGPT d) attribution from other(/future) AI licensees. If your code is scraped once but effectively viewed 1 million times from GPT, you won't see those 1 million hits show up; you can only vaguely infer they might be happening if the attribution is actually implemented, and some users click through on it (or by reverse-querying the AI). So c),d) will proportionately increase as a),b) proportionately decrease.
So everything has changed. And obviously the incentive to you to continue to provide unpaid volunteer labor ongoing without even attribution decreases.
[0]: https://creativecommons.org/licenses/by-sa/4.0/#ref-appropri...
> But if you delete your comments, you won't get anything at all anymore. If they remain, real humans will still benefit directly or indirectly. And why should I write exclusicly for openAI? I share my knowledge for anyone.
The goal -- implicitly for AI companies and explicitly for many of the commenters on this story -- is to replace sites like Stack Exchange. Stack Exchange's traffic will instead go to ChatGPT. The most likely outcome of this is that Stack Exchange will eventually shut down or severely degrade its service. If ChatGPT were a supplemental tool, one user out of many, you would be right. But it's not a complement, it's a competitor, designed to make a profit off of assimilating my work without giving me any compensation or credit.
With ChatGPT and similar platforms, trained on SE answers (and open Github repos,...), people will eventually skip Stack Exchange and directly go to ChatGPT.
At some point, passing the hoops turns into a negative predictor of comment quality.
It doesn't help that asking for a more recent answer gets your question closed as a duplicate, and new answers can never overcome the inertia of the historical ones.
September has came for stack overflow
We've outscaled our original hopes for the Internet. It was originally meant to be a tool genuinely controlled by it's users; unfortunately, it's largely ended up in the stranglehold of a few monopolists.
I have not really had a lot of accuracy issues with GPT, but then again, I probably and not savvy enough to spot them, most of the time, anyway.
Everything we write on the web is like that, including this very comment.
the answers have already been exported. all you're doing by deleting it is ensuring it's only available on ChatGPT, and no longer available to web users who aren't using AI tools that ingested the content before it was deleted.
For years the community has defended punitive down votes on correct answers to crappy questions as "you can do with your vote as you like". I see no argument against flipping that around.
Are those open source llm users in turn selling access to the content they got for free, and also stripped of attribution?
What exactly is changing hands in trade for the money, that doesn't one way or another violate CC-BY-SA?
It's not merely the fact of any form of commercial activity, since there is no NC in there, but the specific actions here by both StackOverflow and OpenAI violate the terms the content was originally created and shared under.
i haven't used that website since GPT came out, and now i contribute nothing to it
but i'm glad all of its content ended up training the models that put it out of business, thanks SO! you'll never be anything other than user contributions
what's amusing to me is that some people even in this thread are calling it a pro, not a con. I guess our field does indeed attract a certain kind of personality.
Now they paid the shakedown fee, and Stack Overflow has user riot on their hands, albeit one that's trivial to shutdown. Send OpenAI the db backup.