StackOverflow is banning accounts that delete answers in protest against OpenAI
m.benui.ca
m.benui.ca
But if anyone has a license to use your content... Then certainly Stack Overflow do! If they want to post an old version of your question/answer that hasn't been overwritten with a message about some cause you care about, important as it may be to you, then they're free to do that. This is not news.
The user didn't delete answers, they updated them to say something unrelated to the question. I understand the protest, but there's a difference between deleting and what's essentially defacing.
This problem is not a technical one; the companies already have the data.
Yep. For instance in this paper[1] they asked an LLM a bunch of political questions to figure out what its political alignment is.
[1] "Discovering Language Model Behaviors with Model-Written Evaluations" https://www-cdn.anthropic.com/e4f69aacd8c0905030172bc6eb480c...
So instead I changed my highest-rated answers to a protest message.
Within an hour mods had changed the questions back and suspended my account for 7 days.
How about repeatedly changing just a few characters, corrupting the post over a period of time? (Entrepreneurs: you can have "SO post gradual corruption as a service" for free, you're welcome)
2. This seems easily detectable, by checking edit volumes of a given user and/or post. Even 1 edit per month is probably going to be suspect compared to how infrequently posts actually get edited.
Detectable, maybe -- month one, a few homonyms which could be spelling corrections: been -> bean etc, month two, (near-) synonyms: bean -> lentil ... throw in a few UFT8 homglyphs to add noise. Unlikely that that sort of corruption would be obvious to a machine, then manual checking is needed, and do they have the manpower for that?
Also one could also use a "masking" edit. Make the defacing change then another more innocuous one.
> throw in a few [UTF8] homglyphs to add noise
That seems like it could be defeated by a pretty straightforward normalization task. Either just a straight mapping or a mapping that's applied contextually (e.g. replace an omicron with an o, but only when the adjacent characters are ASCII).
What about them? Nothing's perfect. One rather effective of was of undermining any kind of effective protest, is to pepper it with demands for perfection that can only be satisfied though inaction.
And I think there's a subtle contradiction in your objection: by doing something like this you're also sabotaging the thing that is slated to replace many of those those "student/junior programmers." So are you really doing them harm?
I find for web technologies it's less useful because the technologies keep change, and what was a good answer in 2010 is not now. To answer questions I find it more useful to consult a primary source, such as the documentation on a some software's website, or a forum dedicated to that software.
On SO I often get no replies, or unhelpful and often toxic comments along the lines of "Technology X is bad, use Y instead".
Or I often get people who answer the question as if they show applying for a job at Google. No, I just need help with some software to process a few thousand rows and send a summary once a month by email. I'm not trying to process petabytes of data with a response time of <1ms.
Seems fair, and provides back the credit to the person who prepared the answer.
Whether LLM training counts as "copying" has yet to be settled in court. Absent that, it's unclear whether attribution is even required. For instance, facts aren't copyrightable, so you can theoretically incorporate facts from other works into your work without attribution.
SO can actually do this?
I was surprised by this partnership to be honest, but I don’t blame contributors picking a side on this. SO is merely a provider of infrastructure, but the way that OpenAI wants to monetise people only benefits them and SO directly.
Not everyone agrees with OpenAI, etc. This is about giving people an option as opposed to be being told “no” and thats the end of it.
You could avoid the scraping by requiring login or maybe having some continuously generated AI noise requiring social context filter out that a scraper wouldn't have (e.g. a bot in a chatroom that spews garbage that everyone knows to put in their killfile).
This is easily bypassed by having some sort of pattern matching sign-up script (because realistically, everyone is going to be using an off-the-shelf solution like wordpresss which you can detect and write scripts for) and paying a captcha solving farm to solve any captchas for you. Practically speaking though, having a signup wall is going to be killer for getting visits. It's already hard enough as it is to get started as an independent writer/blogger with no sign up wall. If you're going to force people to log in as well all that will do is ensure your posts are read by no one.
>or maybe having some continuously generated AI noise requiring social context filter out that a scraper wouldn't have (e.g. a bot in a chatroom that spews garbage that everyone knows to put in their killfile).
Similar problem to the above. Why would someone want to join your community when the first thing they see is that it has a spam problem that's seemingly unaddressed?
> This is easily bypassed by having some sort of pattern matching sign-up script
I would imagine it's not that easy. At a minimum you could have captchas, and if those stop working you could always add some other friction (like a nominal payment for an account, a la Metafilter).
> Similar problem to the above. Why would someone want to join your community when the first thing they see is that it has a spam problem that's seemingly unaddressed?
Because they understand the reason, and there emerges some common understanding that's something you need to do to discourage scraping?
The only way I can see this working out is in some sort of author's forum, where everyone there are personally affected by LLM training and therefore they're willing to put up with the barriers. The same can't said be true for the average reader, or even most readers. They're far less to care about their artistic/literary livelihoods being killed by AI, and as we seen with how popular Amazon and Uber is despite their treatment of their delivery "partners", even if they're sympathetic to you that's no guarantee that they'll support you at their inconvenience. The overwhelming majority of readers is going to hit your signup/pay-wall and bounce, which means all you're doing is kneecapping your reach in exchange for a bit of smugness against AI companies.
Huh? I don't think "smug" is an appropriate word to use there. One isn't smug for not wanting to get exploited.
If anything's smug, it's this sense that AI companies should get whatever they want and everyone just and there's nothing anyone can do about it.
currently? Even in the geocities days, companies can presumably crawl/scrape your site to train their LLMs. The only way of preventing that is by paywalling your site, which I doubt many hobbyist bloggers are going to do.
Edit: https://hachyderm.io/@ben@m.benui.ca/112396508798014598 has it cached for me.
> Here's a cached copy of the linked post on a server with more capacity
Alright, if they use it to train some AI - what is the big deal? As I understand the original posts are still accessable. Essentially users will be paying for faster search and better precision since you can do the same without payed AI.
If they didn’t change something they would simply cease to exist. Does anyone have a better idea?
Also, answers are edotsble by other users, so you don't own them after you post them.
Let's be honest, your answers on StackOverflow was never _not_ used for profit.
And before anyone brings it up: GDPR does not mean "every single bit of information I ever wrote can be deleted whenever I want". Not on HN, and certainly not on Stack Overflow where you agreed to have everything under the CC license. It's painful how badly some people understand the GDPR.
https://www.osnews.com/story/139601/stack-overflow-signs-dea...
From the privacy policy:
> you represent and warrant that you own all right, title and interest in and to such User Content, including, without limitation, all copyrights and rights of publicity contained therein
So I'm pressing X for doubt, but someone with legal knowledge may enlighten me otherwise.
For example, pasting in the text of a paywalled article.
It doesn't mean "you retain ownership of the content."
One might reasonably claim that SO is doing something equivalent to saying "this person thinks the answer as written is true and implicitly endorses the products involved". YMMV, but publishing something like this, especially with the knowledge that the person would like the post removed, could be considered defamatory.
You can delete your account; that's not a problem. I know because I did that myself in spite of having quite high reputation (for the politics site, due to rampant antisemitism there).
After that it will just show as "user12312".
Does OpenAI comply with the CC-BY-SA 4.0 license terms? AFAIK (I don't use their products, so don't know much), they don't actually include 'derived from original by ...' and 'licensed CC-BY-SA 4.0' in their works, and StackOverflow has no rights to relicense the content to a different license.
I'm not defending Stack Overflow and/or OpenAI here; I also pretty much stopped all my contributions for various reasons. I'm just saying it's fine for them to have reasonable protections against sour grape deletions, and that GDPR has nothing to say about this.
Well played, well played.
Strange game: https://www.youtube.com/watch?v=uOoXwxqeVzg
That's how the internet worked since web 2.0. Generative AI training is just another way it's being monetized.
I doubt they had any reasonable expectation that Stack Overflow was a charity.
Also my answer about centering a Div… really isn’t worth much.
There is absolutely no real issue unless one starts arguing about slippery slopes where Stackoverflow monetizes its platform and the website becomes unusable due to AI spam etc. I find these fears quite unrealistic.
2, 3 years later it's no longer hanging their bedroom, and you ask, and they tell you they sold it for $50k.
You'd be a bit annoyed on multiple fronts....
The analogy fails because in this case, the painting is still on their bedroom, just like how the guy's stackoverflow answer is still on stackoverflow. Moreover, when you contribute to stackoverflow, you agree that your work is licensed under a creative commons license. So it's the equivalent of giving your friend a painting, telling them "oh btw you can do whatever you want with it, as long as you attribute back to me", and then get mad that they digitized it and put it on their blog. Finally, let's be real here. This isn't a painting worth $50k. No one's comment is worth anywhere near that.
It's hilarious how many of the same talking points from discussions on piracy from a decade ago has been reused in discussions about LLM training data (eg. "piracy = theft" -> "LLM training = theft")
So long as OpenAI abides by these terms, I don't have any problem with them using my answers for training. Do you think they will?
Do not submit answers to StackOverflow, write articles for your own blog/personal web-site instead.
i would say this is the so-far elusive web3.0
Editing the answers through SO's tools is not the right way to protest (since it's trivial for them to suspend you because you're clearly violating their rules); the GDPR is. Making a GDPR deletion request is quite easy.
No one is editing their posts constantly and with high change.
Take your backups (you have anyway), diff the content, allow only minor changes and a max edit limit of x, restore the content if those criteria are not met.
Sure it costs a week or a month to add this feature to the platform but data will not be lost.
Nonetheless, SO is not the only ones having an archive of SO
Any protest like that will affect SO's ability to do other things they care about.
Protest is rarely pretty, often not legal. That's the point...
Lawyers are a little more expensive than handling a click, especially at scale.
And GDPR is an eu thing not a usa thing.
I don't think it's been tried before but that sounds like a possible legal path to taking down the content
In the case of open source software, it could be some use which reflects badly on yourself.
Here in this scenario, the angle could be that using it for training in AI could affect the integrity of the content produced
In the case of StackOverflow post are a contribution to the collective knowledge base rather than personal data. If your posts do not contain personally identifiable information, StackOverflow can easily argue that the content does not fall under GDPR’s scope for deletion based solely on personal data grounds.
Surprise your content is publicly available and used by a company for profit \o/
Srsly... your QA was used in GPT before this agreement already now SO profits from it which makes it in theory better because now SO (the platform you liked before this) is able to continue investing into it and making money / profit which motivates them to keep investing and inovating.
Why tf is this a bad thing?
Have you never heard of copyright or licenses? Or what is your point?
SO is not GitHub. GitHub has high motivation to protect your (c) for you, SO has not. Never has.
GitHub already sells Copilot, trained on copyrighted user data. Where do you observe this "motivation"?
Just shows how much money OpenAI has to spend when they could have got it for free.
... But the easiest way to shortcut all of those thorny open questions is to go through the front door with a big sack of cash and an explicit agreement on how the data will be harvested and what it will be used for.
I thought it was ruled to be fair use? eg.
https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....