Moderation strike
meta.stackexchange.com
meta.stackexchange.com
https://openletter.mousetail.nl/
The meta post linked here is targeted more towards an internal audience of active users.
There are two big parts to this issue, one is that the company is overriding the decisions of the communities and essentially preventing them from moderating AI-authored content entirely. The second one is the way this was done, with no feedback at all, extremely quickly, with vast differences between the public policy and what they told the moderators.
Unfortunately, this seems a naive take; the core mission of the network is to serve the commercial purposes of the business.
We may have been duped by a company's lies, of course. Seems like they've committed a mix of copyright infringement and fraud in that case.
Expecting a business - with infamous episodes of contempt towards moderators - to behave in the idealistic way presented here is naïve.
Might it submit to pressure? Perhaps. But the wording of this letter in its presumption of the motives of a private organisation misunderstands the reality of the commercial world.
Are you familiar with strike tactics?
But that's getting off the topic of whether or not the wording of the quoted paragraph reflects reality.
Not in this case. The exchange is real but non-monetary.
>But that's getting off the topic
No, not really.
not off-topic at all - but I think you just noticed that your argument doesn't hold.
That's an incredibly naive take. Do you think there is an abundance of moderators just waiting to fill the ranks? Several SE sites are struggling with a lack of moderators already.
Similarly, there is no abundance of community members curating questions and answers. It's all volunteers, and it's not like there's a huge untapped pool that they can access on-demand.
It's kinda like saying that the entire community can be easily replaced because it's free. No, the community is the business value!
Lastly, the main anti-spam tool of the site (smokedetector-se.org - community-built and -hosted) is also offline as part of the strike. The tooling built by SE the site is in no way shape or form adequate for combating the amount of spam the site receives. Sure, it's not irreplaceable, but it's not "free".
Hence recent developments of large brands committing billion-dollar acts of seppuku in near-realtime: somebody is pulling strings in non-obvious, non-"free market" ways.
It's a zany time to be alive, and one hopes that sites like this one can help redistribute a "free market" sensibility that seems on the wane.
Could you cite examples? I don’t really keep up to date with the news so I would like to know?
Is this the index fund conspiracy theory again? Where companies like Vanguard supposedly form some sort of shadowy evil cabal that secretly controls the entire economy?
Specifically, organization of contributors actually sounds like a very effective way of holding these morally dubious “content” middle men accountable. The consumers are too heterogeneous, numerous, and uninvested to realistically coordinate.
When companies talk about their "mission", they are referring more to how they intend to deliver value to their owners, usually by identifying some social need and satisfying that need in exchange for value.
If the latter is done based on a lie, then you cant align your “commercial purpose” and people who fulfill it..
That's exactly what moderation is for, checking, questioning, etc, its not an argument against AI-generated content.
There seems to be an emotional reaction, but people will still value the highest quality content, whatever its provenance.
If the argument is that the workload will be too high, well make that argument - don't sidetrack with misleading and idealistic appeals to emotion.
The reason we created "corporations" as a concept was not simply to have a vehicle to make as much money as possible, any way possible. It was to serve society by providing services and making products, which would then be sold, and, if they were good enough, would make the company a profit.
This idea that no business should ever be expected to do anything but what makes them the most money the fastest is toxic and is ruining our economy and our society.
Workers taking back some of that control is a great thing.
The business is a helpful abstraction layered over the top of all these people doing the actual work. It's useful as long as it keeps the lights on. When it stops doing that the community has the right and the responsibility to move the work product of the huge quantity of people actually doing the thing here somewhere else.
> Stack Overflow, Inc. has decreed a near-total prohibition on moderating AI-generated content
That would have been a really good place for [one of these](https://somewhere)
The public policy by SE is misleading, it makes the rule appear a lot different than it actually is. I am a mod on a small SE site, so I have seen the internal communication and it does essentially prohibit moderating AI-generated posts except in some very narrow circumstances.
If a post isn't constructive, they can and should moderate it away, whether or not it was AI generated. If they want to have a rate limit on number of answers per day, they can do that. But they need to moderate based on observable metrics, not guesses at what's happening behind the scenes.
As a person that looks at stackoverflow answers, it's also awful to see the long-winded AI answers. If there's mistakes in the answers, there's no way to get the submitter to make changes because they likely didn't understand it in the first place.
There's ways to do it correctly, but inexperienced developers opening chatgpt and blindly copying answers to harvest karma isn't it.
Banning all moderation without consulting the community is not a step towards coexistence, it's a step towards turning the site into a dumping ground for stale LLM spew in the name of engagement.
Companies like Stack Overflow seem to care less about the quality and accuracy of the site's content as long as the engagement numbers look good. Unfortunately, a lot of users don't seem to care about quality, either, which is how we've wound up with the current state of affairs where just about every social platform is increasingly flooded with bot spam, misinformation, scams and junk.
By banning low-quality posts, including unverified LLM answers, the site can secure a future for itself as a bastion of quality. Or it can turn its back on the experts that built the content that's used to train the LLMs and hope that LLM quality improves enough that expert humans aren't necessary.
Even if that gamble works out for them and LLMs do mostly replace humans as some commenters optimistically seem to expect, then there'll still be no need for Stack Overflow, as one can simply ask an LLM directly. That's why effectively dumping human experts for LLMs seems like a failed business model either way. The best approach to move forward seems to be to carve out a space that provides unique value that LLMs can't and ride that until LLMs make significant advances.
Disclosure/context: I'm a daily SO answerer on strike, among the top ~4k users by rep overall. I'm pretty sure most folks who are dismissing the problem don't monitor tag feeds or curate queues enough to see the flood of blatantly wrong LLM answers spamming in from new accounts.
A user has been writing (bad) answers in broken English for months, and now suddenly writes answers in perfect GPT-esque style while the technical aspect is still bad.
The moderation burden for the second kind is much higher, while other users are much more likely to mistake it for a good answer. If the policy is "no AI-generated content allowed", should moderators be allowed to suspend the user?
Careful, clear writing is hard even for native speakers, and good writing serves as a proof-of-work: answers that have had more work put into them are a signal that more work has been done, i.e higher quality answers.
We're quickly losing that signal.
The new policy is that you have to follow all the normal policies for all the posts, you don’t get to pull out the banhammer just for a suspicion of AI. And the moderators are striking because their want to keep the power to issue uncontestable 30-day bans whenever they feel like it?
The result is pretty much that AI-generated content is essentially allowed as it cannot be effectively moderated. Even though many sites still have an official policy that disallows it.
Disclaimer: I'm a mod on a small SE site, though I have not acted as a mod on any AI-generated content.
If mods could decree a piece of content as AI generated (and delete it) willy nilly, then that would be far worse IMO.
“helo plz can help w cod, is broke has error”
And then the next day posts four answers in perfect English to four different topics, with that GPT “vibe”.
You can’t reliably detect generated content in a vacuum, but Stack Overflow is a very metadata-rich environment for moderators.
Why censor good answers?
It is _effort_ to moderate content, to read answers and see if they are legitimate. It is almost zero effort to chatgpt out "answers".
The people providing the ChatGPT answers _do not care_ about them, they are only looking to pad their Rep.
It is an attack against the very core of the StackExchange system.
Also, theoretically (never happened), if someone mentioned their exceptional SE profile on their resume or their cover letter, I'd for sure ask them some details from the topics they suggest to be experts on.
Will they also bring ChatGPT to the interview? That'd be fun to watch.
It is already done, as I've heard. The previous iteration was to pay your lookalike to pass the interview for you. Or just pay anyone and blame lack of camera on "technical difficulties".
If you are an applicant to a position where there are 5000 resumes submitted and the people doing the filter want a quick and easy numerical ranking of them - provide your Stack Overflow account.
At that point, they can look up your rep and pick the best ones based on that.
This doesn't happen as much in US based companies as there are other metrics that they can use to filter candidates rather than SO rep.
However, if you are in India and applying to a consultancy and every resume and transcript are very similar to the point of not being able to distinguish between them - SO rep provides a very easy way to rank and filter applicants.
Unfortunately, I can't verify this. I've only heard it second hand but it does make sense and helps explain why I occasionally get SO account and stats on resumes from contractors.
Because a bunch of dumbass software companies (especially the sweatshop ones) decided that your annual review now needs to include a bunch of dumbass "open source software social work" or you get dinged on your review.
So, Github contributions, Stack Overflow moderation, etc. all became subject to Goodhart's Law: "When a measure becomes a target, it ceases to be a good measure."
ChatGPT is just accelerating this contamination of the well to all-out flooding of it with industrial sewage.
Stack Exchange could just call an LLM API when a question is asked and show the response. That wouldn’t be near as valuable as their verified index of answers.
So if e.g. a user posts a dozen long answers within 10 minutes, and they all have characteristics of ChatGPT, that would be a pretty good signal.
Ah yeah, the frequency of posts definitely makes sense.
> output of specific AI tools like ChatGPT produces certain patterns that are quite noticeable
I would agree with you for gpt 3.5, but I don't think this is the case for GPT-4 (I've spent several hundred hours using GPT-4 for various tasks - mainly related to coding & learning random subjects).
Say, a person has an answer but English is not their native language and they manage to stir ChatGPT into writing a good answer. Would we prefer to have that posted instead of keeping a question hanging without an answer at all?
The only issue I can see with AI use is the rate of new content generation. Recent models are quite OK at giving a decent answer. SO is not a pinnacle of exceptionally well thought out answers from people either. There are great detailed and well sourced answers but more often then not you get an incomplete, outdated or even just plain wrong answers. Bespoke artisanal hand-crafted ethically sourced answers from fully organic free range humans that still lead to stack overflows and misalignments elements on webpages.
In practice, as reported in comments at https://meta.superuser.com/q/15021/38062 and in many other Meta Q&As, the answers that people are lazily machine-generating at high volume are far from correct; and the consequent upvotes that they garner reveal the unsurprising fact that there are a lot of people who vote in favour of things based upon writing style alone.
A separate question is why there's still a lot of crap questions/answers on SO if quality is the goal? There's a plenty of low-effort and incorrect answers made by real people that are not penalised in any way.
Why? It becomes irrelevant if an account is spamming, sure delete everything regardless, but if I have a generated and correct answer in my otherwise pristine account?
What I'm trying to say is, the fact that some people use it to spam shouldn't make it a simple ban condition. Otherwise that'd be banning emails to fight spam.
You haven't. People aren't. This is a hypothetical that isn't the reality, and an irrelevant distraction.
Go and read the comments where I just hyperlinked, then read the months of back-discussion on this in the other Meta Q&As that I mentioned, starting with the likes of https://meta.superuser.com/q/14847/38062 right there on the same Meta site, and a lot more besides on many of the 180 other Stack Exchange sites, continuing with the likes of https://math.meta.stackexchange.com/q/35651/13638.
How can you be so sure? If one day comes a 100% reliable way to detect all AI-generated responses, how can you be sure that also the good ones won't get deleted in one major sweep?
Yes, I see there are many people who despise the AI generated spam on many sites. But nothing you posted proves that all (I'd even say, "significant portion of") AI generated content is spam.
I don't see anything wrong letting the AI generate an answer and edit the rough/wrong parts if necessary.
And this is not what the previous moderation policy was trying to prevent. What it was trying to prevent is answers from people skip that second step.
> "But what if they were correct answers?" is largely an irrelevant hypothetical side-issue.
Only if you don't prompt it to do otherwise.
i'm really not seeing the problem here.
What is the point of a software Q&A site anyway, why not just read the docs?
In my experience, Q&A is useful because it summarises and consolidates disparate information into a concise response to a prompt (sorry, I mean to a Question!) which is... exactly what a chat AI does.
Isn’t chat AI an existential threat to SO, even if AI is banned there?
Certainly I find chat AI better in many cases, and if I don’t like the answers, just going to the docs is the next step, not a generic Google search with SO in the SERP.
There is knowledge you will never find in documentation. 50% of places where SO has been helpful to me are not "consolidating documentation". They are solutions to obscure bugs, Useful APIs which are not documented with any example, and low-level logic / high-level design solutions.
The offence here is spamming, AI is orthogonal. Users spamming should be given a timeout.
You can make broad classifications for an act (it wasn't murder, it was negligent discharge of a firearm), but that doesn't preclude the fact that it is also something else (shooting in the direction of a person without regard to their life and killing them is murder).
In this scenario, spamming with or without AI would carry the same penalty. The benefit is that you don't need to determine if AI was involved, since it doesn't matter.
Similarly, use of LLMs to generate and post large quantities of text from a short prompt is inherently spamming. The user could provide exactly the same value by posting their prompt. If the prompt isn't valuable, the output won't be either, as anyone with a copy of the prompt could generate equivalent output for themselves if they wished to.
How is this different than people posting things that they didn't test themselves?
try
{code block}
Hope this helps
Answers that are clearly not an answer are edited to be pretty rather than down voted and flagged because they're wrong ( https://stackoverflow.com/a/76402243 ).The problem isn't an LLM (though that just provides more scale) - its that incorrect information isn't removable / actionable on SO. If the person tried to answer a question it remains up.
There are definitely alternatives to the unilateral ban (Thinking about how, like, chess.com does bans based on people cheating), but saying "AI content is qualitatively no different" ignores the bigger ecosystem problem from having the average answer quality simply go down
I certainly don't have the answers (and let's face it it's likely going to be a problem here too).
So, like the current organisation of things, that works quite well? You put entirely too much value on moderators being perfect. Moderators have been biased forever, have banned for no reason forever. It's a website, you can live with that (or without it). If they're unreasonable, move away, find another place or create another place. If you can't move away, deal with it, create an alt, don't get caught doing the same thing that got you banned, and that'll be it.
This is true, and it's why the policy promulgated by the moderators was explicitly temporary. It was a response to an acute problem to give everyone time to figure out how to handle it as a chronic condition: https://meta.stackoverflow.com/questions/421831/temporary-po...
LLMs are great at producing bullshit that looks convincing.
Essentially, one has to trust responders to provide answers to some extent. Untrustworthy responders can use generators to bullshit their way into acquiring trust.
> hoping
Hmmmm
The community of moderators is kind of a symbiote attached to this enterprise. It gleans and curates and makes the end product more helpful to users. "Helpfulness" is a second-order effect of this moderation, and the whole attraction to the business.
After telling moderators to not moderate, moderators should get the message: It's not about AI, its about whether moderation is valuable, and whether helpfulness of answers is valued by the business.
Ever since the posting system was turned into a social score where people are mostly conncerned with increassing their score versus answering questions, stackoverflow has failed it's users.
Just another example in the very long list of for-profit plaforms doing what's best for profit over what's best for the users...
Then the second point is, on a blind sample of question answers, how well can they tell if something has been generated? I bet it's not stellar.
I hope Stack Exchange stays on their position
I also don't see the end goal making a bot that answers questions on SE. It doesn't make money. Maybe to get points? But once you're past their thresholds there's no reason to keep doing it, and you get there quick. Accounts don't get sold to advertisers like on reddit. You'd at most do it once. And who would even do that? The very narrow niche of people who'd want to boost their SE points? Maybe you're shooting for the leaderboard but if that's the case... you'd get noticed, wouldn't you?
I could see some kids doing it, if that's the big threat they're facing... then they're over-reacting and it still doesn't warrant blanket bans on LLMs
A tip I learned long ago: Never ask a geek "why?", just nod your head and back away slowly.
What a GPT ban does do is give moderators an extremely subjective tool for removing answers. The GPT detectors are unreliable, so that leaves the moderators doing a gut check on whether this particular case is a false positive.
Instead, why don't they just use the GPT detectors to detect answers that might be wrong and then moderate based on existing policies?
Now spread that one bot over hundreds of users and you have the same end result. That's why the communities of different SE sites all ended up with a similar policy.
I'm not suggesting either problem is easy, but taking the "ban AI" approach is harmful for several reasons. Detecting AI without false positives is extremely hard and puts too much decision-making power in the gut reactions of moderators. Additionally, banning it has the negative side effect of making the tool unavailable to people who might legitimately benefit from it when producing good content, such as ESL speakers using it to polish their English.
We need to moderate based on the effect of the user's content on the community, not the technology used to produce the content. If you're dealing with a bot posting 10k questions and answers per hour, it wouldn't matter if that bot were using GPT or just re-posting content scraped from the web—the abuse isn't in the use of AI, it's in the spam. So make a rule against spam and auto-ban people who do it.
And, I guess they will just realise it is AI generated when someone notices that what is written does not make any sense, but it is still really well written. In any case, very difficult to identify, for sure.
- they're all written in the same style ("It looks like you're trying to do convert a string to an integer...") - they have perfect grammar - they're often/usually wrong - they tend to be tone-deaf to the question or are answering a different question than asked - their formatting is often poor by virtue of copy-paste from ChatGPT and the author's lack of familiarity with markdown - they're usually posted by new or new-ish users that have few reputation points (ostensibly, their goal is to farm reputation rather than help curate a quality resource) - they're posted in large quantities (a dozen or more in an hour) that'd be virtually impossible for a human to churn out (especially a brand new user!) - they're posted in random tags that show no clear connection from one to the next--most answerers are subject matter enthusiasts or experts and stick to a narrow range of tags - they tend to contain hallucinations obvious to SMEs, like invoking methods that don't exist - the code is plain wrong when executed - the code style is "boring"/"vanilla" and tends to steer clear of idiomatic language features that a seasoned programmer would employ and has few formatting quirks that a human might use - the code is often heavily commented in a predictable and artificial manner - the explanation after the code and overall layout/flow of the post is often the same - they tend to be dramatically different than the user's normal answers, which have typos.
As you'd expect, it's difficult to _prove_ that a particular answer was generated by an LLM (the fact that LLMs can't reliably detect themselves is part of the problem--they're inaccurate!). However, the possibility of occasional false positives (SE has provided no actual evidence of this being an issue), seems a necessary price to pay, and could be solved in a more balanced manner than prohibiting all LLM moderation. SO would be unusable if it became a stale cache of mostly-incorrect LLM spew, which is what SE's new moderator guidelines seem to be OK with.
If I want an LLM answer, I'll ask an LLM. I can then prompt engineer and iterate. If I want a human answer, I want to be able to ask on SO so I can be guaranteed I'm "speaking with a human".
Disclosure/context: I'm a power-user on strike.
You can't treat volunteers like that. You have to keep them happy. You have to treat them with respect. You're not paying them for their free labour.
I don't like it on Reddit and I don't like it on SO. They should pay professional moderators just like Twitter and Facebook.
Which are?
I understand very well that it is active moderation which keeps SO from being a toilet of spam and junk. At the same time, in my limited experience SO moderators are petty tyrants, very jealously guarding their tiny domains.
I often wonder about this. I can see how this is true for somewhere like Reddit where lazy memes could overrun any subreddit, but I'm less convinced when it comes to SO. The vast majority of SO moderator actions are on questions that non-mods have already downvoted and reported. These questions could/would be hidden by SO's algorithm anyway. Only new queue lurkers see unpopular questions.
I think that SO's moderation could be almost entirely automated and overseen by a small number of professional paid moderators.
I don't think that's always the case. People can and do organise themselves in ways that prioritise collaboration over treating these sorts of things as a zero-sum game. Not all orgs have to function in this way.
If this is all you have experienced, then I urge you to look beyond your current circle (or bubble) to find people who genuinely care about others and want to improve the world.
Since most people tend to avoid confrontation, that gives the kinds of people who want to take as much as they can (or to have their way in some other way) a free pass until someone does in fact fight back.
Hence it's not surprising if, in the big picture, it seems like "people" (as in enough of them to make it a potential issue) do in fact seem to behave that way.
> Content posted without innate domain understanding, but written in a “smart” way, is dangerous to the integrity of the Stack Exchange network’s goal: To be a repository of high-quality question and answer content.
Tricking our bullshit detectors (eg cues present in the text or implied context) is the biggest problem I have with usefulness of AI generated content.
I wrote the Medieval Content Farm (https://tidings.potato.horse/about) as an excuse to talk about this (and cope), although people focus more on the fact that I present it as a joke.
I don't see the end goal making a bot that answers questions on SE. It doesn't make money. Maybe to get points, but once you're past every point threshold there's no reason to keep doing it, and it happens fairly quick. Accounts don't get sold to advertisers like on reddit. You'd at most do it once, and for the very narrow niche of people who'd want to boost their SE points? Maybe you're shooting for the leaderboard but if that's the case... you'd get noticed, wouldn't you?
I could see some kids doing it, if that's the big threat they're facing then it doesn't warrant a blanket ban on LLMs. I care about 1) spam, 2) answer usefulness. Not how the answer was written
Computer generated answers compiled from existing resources may work for simple questions, but not for things that require specific knowledge and experience.
I still get a better hit rate from mailing lists and github issues.
But since then 1) moderation got more severe, so you can't ask a question like "what packages are there for X?" (even though there are many remaining from the early 2010s) 2) many questions similar to older ones but different in fine details, get closed quickly.
I see new tech go to their own forums based on Discourse.
So, since 2020, I'd still come to SO for answers already available, but for a place to make new questions -- I'd look elsewhere.
The SO problem isn't AI, it's people submitting low value answers, regardless of the way they used to produce those.
It is if anything a failiure of their reputation system, both in incentives and in repercussions.
Just because you are an expert does not mean you tackle every mundane but needed part of the work bare fisted or you memorized every edge case.
It saves me hours working on mundane shit every week.
Similar situation with moderation took hold in Wikipedia back in the 2000s, when it became only up to them whether a paragraph is "neutral" or "an opinion" and must be deleted (e.g. some pages have pieces saying "it's a common misconception that ___", but in some pages they got deleted with edit comment "it's a POV, must not be in Wikipedia").
> Here is how platforms die: first, they are good to their users; then they abuse their users to make things better for their business customers; finally, they abuse those business customers to claw back all the value for themselves. Then, they die.
[1] https://pluralistic.net/2023/01/21/potemkin-ai/#hey-guys
This is a boycott, not a strike.
However, as a counterpoint, I now see people who spend a lot of time using ChatGPT actually end up writing in real-time, in-person, like ChatGPT's default vanilla output. Just like some American kids now say "mummy" because of watching so much Peppa Pig.
In general no. But I don't think anyone human writes in the typical style of a ChatGPT answer, so in practice there is a large class of cases where you can tell.
The ability to determine the origin of content is indeed a complex issue. However, it's important to note that researchers and developers are actively working on developing methods to identify AI-generated content. There are ongoing efforts in the field of AI ethics and responsible AI development to address this problem.
Some approaches to tackling this issue include developing AI algorithms that can generate AI-generated content, also known as "adversarial AI." These algorithms aim to create AI models that can detect and identify AI-generated text by analyzing patterns, linguistic cues, or statistical properties specific to AI-generated content.
Additionally, researchers are exploring the use of metadata, watermarking, or digital signatures to provide additional information about the origin or authenticity of content. These methods could potentially help in distinguishing between AI-generated and human-generated content.
While it may be a complex problem to solve entirely, the development of techniques and tools to identify AI-generated content is an ongoing area of research. The aim is to create a balance between the benefits of AI-generated content and the need for transparency and trust in online information.
Of course, all paragraphs above are ChatGPT-generated. I did not even read them before commenting. That's the bar to clear. Otherwise, see https://xkcd.com/810/
I see one possibility for them, embrace AI and generate a quick/automated first reply by AI marked as such (with a disclaimer) for every post. It should be subject to the same voting system.
The error of moderators and SO here is to discard AI generated answers because some wrong (but sounding right) answers it can generate... when often AI answers are also correct and even at times out do what a single human would have found/answered.
If you can harness the existing human knowledge and correct the "bad" ones (badly rated AI answer given less weight) to feed the models of the future it seems like a win for everyone in the long run. A ban misses this opportunity and generates even more work for moderators which will inevitably also ban some innocent users and valuable content.
Problem is, most contributors would probably stop contributing (answering questions) if that were the case. If there's an automatic answer that is correct 2/3 of all times, that would mean lots of time spent reviewing automatic answers and lots of time "wasted" (where a contribution isn't needed), which will probably discourage most of them
You are an optimist
As for the correctness of the solution: that's why the voting system and tickmarks are there. Wrong solutions would ideally be downvoted and never marked correct, I don't really see how AI is making a big difference here. Moderators today aren't running the answered code either.
But why should SO exist for this? If people wanted an LLM to answer their questions, they could just go and have that. There is no purpose in caching stale LLM answers on SO.
Which means future OpenAI models would be getting trained on the output of competitor models. Like Bard. Oof.
One of the more exciting scenarios would be if it turns out that performance can improve indefinitely with a generate -> filter -> train cycle. There are certainly parallels to how humans learn.
I see this said a lot, but in reality it just depends on how the network is trained and how it's prompted. For example, you don't get dumber because you read children's books, you just get better at understanding what makes a good children's book. It's only if reading children's books comes at the cost of reading other content that you might be dumber as a result.
Similarly, an AI doesn't automatically get dumber because it encounters dumb content. It's only if you're training exclusively on dumb content that it doesn't know what quality looks like that you'd have problems.
Broad training sets (ideally pruned of as much junk as possible) and RLHF in theory should condition the network to reproduce quality content and not simply the lowest common denominator of what's found on the internet.
And assuming all that fails there's nothing stopping researchers from just using past datasets with improved architecture going forward. I mean you'd have to wonder why on Earth OpenAI would even release GPT-5 if it's worse than GPT-4...
There's just no scenario here in which what you're saying would actually play out in reality. One way or another companies will ensure the next iteration of their LLMs are better than their previous.
I don't think this is an accurate analogy. It's not books for children -- well-written material at a lower educational/cognitive level -- it's more like books by children -- which is necessarily lacking skill and background context (connection to reality). Think about how children constantly pass around misleading, invented, and incorrect stories among themselves -- and they don't do it maliciously, they just don't know any better. Legends like "Candyman"/"Bloody Mary" for example.
They need an outside influence, an adult or a book or website, to nudge them out of that knowledge rut.
(Of course the same thing can happen with a (closed) group of adults, too, but it's more of a "natural state" with children because they simply haven't had time to encounter as much knowledge.)
It’s a bit like public key cryptography — if generating text is using sign() to sign a prompt with your model then is there a publicly available verify() that verifies the output came from the model but which doesn’t leak the private model itself?
What sort of things exist like this (other than private, API models recording all text they’ve generated)?
I think OpenAI sells a subscription to a ChatGPT detector. So you can pay for h. Does it work correctly? I don't know. Selling the poison and the antidote seems like a good business move though (I know, you asked "other than private API" - not sure they record the generated text, not sure they don't).
Now, a ChatGPT-generated text (for current versions of ChatGPT) is more or less recognizable so for moderation purpose I would guess you don't really need h, you can smell the bullshit. It has a specific way to be overconfident and it feels like it's giving you a lesson in a specific impersonal way without emotions. Something like this. I have the same kind of feeling when reading a WikiHow page, WikiHow has a very specific and recognizable style to explain things.
I guess you can recognize patterns / specific behaviors on accounts posting ChatGPT texts too, which can help for particularly short texts.
I’m honestly sick of being told I can’t say “Thank You” at the end of my post or other dumb crap these mods waste my time with.
It is worth bearing in mind that if we, "good people", complain about moderation, we only see the parts of it that touch our "good posting". There might be plenty of good that moderators are also doing, which only the bad guys see.
Why is this everyone's default response to this kind of stuff?
If anything the network might have lost traffic because moderators have been too efficient in closing duplicate questions before they got useful answers.
That said, AI garbage posts do have to be fought with fire.
But if the user has has the necessary expertise to make sure that what ChatGPT generated is actually correct before posting it, is there really an issue? It would save them some time and allow more questions to be answered.
Also, to your point about humor, in my experience GPTs are very bad at it. If a post is funny it’s most likely not AI generated. Your expectation that all talking meat should maintain a consistently somber decorum while online is, needless to say, unrealistic.
> Your expectation that all talking meat should maintain a consistently somber decorum while online is, needless to say, unrealistic.
Everywhere on the Internet, probably not. In Stack Overflow answers, it wouldn’t be half bad. It’s what they are for. But I wouldn’t even go that far: for example, another’s answer joke that "That's a very complicated operator, so even ISO/IEC JTC1 (Joint Technical Committee 1) placed its description in two different parts of the C++ Standard." is fine in my book, as that answer is otherwise pretty informative. Unlike the one I linked to before, which is just a confusing mess of random access humor (superfluous Xkcd: <https://xkcd.com/1210/>).
> For larger numbers, C++20 introduces some more advanced looping features. First to catch i we can build an inverse loop-de-loop and deflect it onto the std::ostream. However, the speed of i is implementation-defined, so we can use the new C++20 speed operator <<i<< to speed it up. We must also catch it by building wall, if we don't, i leaves the scope and de referencing it causes undefined behavior. To specify the separator, we can use:
Do you feel informed by this? Do you think a newbie would be? What kind of enlightening insight flows from this? This about as funny as the output of a Markov chain: extremely hilarious… for about 15 minutes, after which it just becomes boring.
Future language models are going to provide better quality answers. At some point the line between GPT and expert will get so blurry, enforcing an AI ban will be a fools errand. Sites like SO will still have value as a cache for the most expensive models.
Almost any for-profit platform is doomed to become a dumpster, in the end. After Joel Spolsky and Jeff Atwood, the original founders, sold it to Prosus 2 years ago it has been free falling.
Stack Overflow should have a built-in AI responder, marked as such, that gives an instant unverified first response, which can then be checked and corrected by human moderators.
If you want a ChatGPT answer to a question, go ask ChatGPT (directly, or through one of the many more focused frontends people have built for it). But Stack Overflow should encourage answers from real humans.
From the standpoint of business strategy Stack is caught in a bind.
Banning bot responses seems like a great idea. Bot responses water down the quality of content, and they risk becoming a library of bot responses.
Allowing bot responses also has benefits. The author of post mentions the false-positives on their detection algorithms. A false positive that leads to a ban or other mod action will piss off real users and hurt engagement. Bot responses may not be the same quality as human experience but it may also be a good starting point to drive engagement.
Stack's best move may be to be more transparent in HOW they are making decisions and share their metrics for policy success with their users. Radical openness may be the only way for a radically open knowledge sharing platform to survive the bot-war.
What is the incentive?
What's the incentive for dang to moderate HN? If it's garbage and uncurated, people probably won't use it because the filtering tools to parse the data are non-existent.
People come for an answer to their question, perhaps they answer the questions of others. Just like how any "community" works.
"Their treasure was knowledge."
https://stackoverflow.com/questions/
Just that typical questions about popular technologies.
Isn't Dang being paid to moderate HN?
I would assume so, yes, but the point still stands: he needs to moderate it or people won't find value in it, thus he'll be out of a position.
What's the value in Wikipedia? The curated knowledge. You only have to go on the average Talk tab for an article on Wikipedia to realize how hard-won most content is, and no, the best outcome doesn't always happen, and plenty of genuine nonsense still makes its way into articles and stays there for years on end.
I used to use LinuxQuestions, a forum site, but found that SE provided better answers for me (better format, less searching), so in return I answered questions there if I could. I also used to blog tricky problems I'd solved on my *nix boxen; SE was marginally easier than making a blog post and would get better responses/corrections.
That was the initial motivation: community participation, quid pro quo.
Keeping SE sites like U&L and AU useful is my on-going motivation, in part I enjoy helping others, in part I enjoy using my Linux knowledge acquired over many years as a user.
Yes, the company get value but all answers are "open source" and it serves the public good too.
It wouldn't be that hard from someone to leak it? If all moderators across different sites can see it, it's not exactly a state secret.
CodeChef gives you an option of AI advice if your submission fails to build. It's very effective at solving beginner problems. SE should do the same.
As a positive side effect, it would make spamming with AI-generated answers ineffective and thus save mods time on answers as well.
But the best SE answers actually convey real understanding to the reader; they go beyond the brief provided by the question, and explain a subject with conciseness and lucidity. Nobody could mistake such an answer for ChatGPT output.
[I have a suspicion that such really good answers may be mostly several years old; I haven't ever thought to try and quantify that]
The strike is not the last last-resort effort. Stack Exchange should not only open source all answers, but also open the platform with ActivityPub. Then, those moderators can create another frontend where they ban the AI users. Otherwise, the moderators will create their own platform without Stack Exchange being the main hub.
It's wild that moderator conflicts happen on Stack Exchange and Reddit at the same time.
Or if there is an AI hallucinated sludge, I want it clearly marked as such
Never disallow something you cannot enforce. Doing so causes people to have to make the choice of losing or being a liar. This is in part how cycling became such a mess, when there was an era where either you do EPO and lie, or retire. Because EPO was not allowed but could not yet be detected.
Ever since the posting system was turned into a social score where people are mostly conncerned with increassing their score versus answering questions, stackoverflow has failed it's users.
Just another example in the very long list of for-profit plaforms doing what's best for profit over what's best for the users...
It sounds like the author's main problem with AI answers is that they are purely textual, and the AI is unable to verify them. Does it mean that the stance will change if the AI bot is able to run its code before submitting the answer? Aren't there already AI agents available that can do just that?
Stack Exchange is not about answering people's questions. There can exist a different site for that and for a site like that it could make sense for current LLM to help.
I'm not so sure about that. I think we are quite close to an AI being able to put two and two together to create something marginally new. Maybe not an LLM by itself, but with some high-level iterative/recursive process like this: https://arxiv.org/pdf/2305.10601.pdf.
What I'm saying is that no-AI policy maybe made sense in late 2022, but will not necessarily hold in late 2023.
https://meta.wikimedia.org/wiki/Wikiask/Introduction
It never got anywhere. If it ever does, I would consider being a contributor.
Moderating based on non-constructive comments makes much more sense—if someone posted 30 answers in one hour and none of them answer the question, moderate that. In practice nothing changes, you're still banning people who are abusing AI, but you're doing so based on a concrete quality metric rather than a flawed algorithm plus gut check.
Obligatory xkcd: https://xkcd.com/810/
I don't like this tone, in the sense that suggesting what ChatGPT does is "simply" and "regurgitates" feels like a biased interpretation of what the tech is doing.
I think it is fair to say: ChatGPT is very good at creating convincing appearing content that is also incorrect. Validating content that appears high-quality in form but is actually low-quality in content is a major challenge for moderators. Banning suspicious accounts that moderators believe are spamming ChatGPT based responses is easier than individually validating each and every post from such accounts manually. As the number of posts from ChatGPT backed accounts multiplies this would become increasingly difficult and time consuming.
I understand their pain and I hope they find a solution. But my gut tells me that whatever solution they come up with will be forced to tolerate some amount of GPT generated content.
This data will all be fed right back into GPTX at some point leaving an even worse version of ai to generate future data to ingest again, ad infinitum.
And if so, is this end-stage capitalism? Surely unpaid workers effectively unionizing because they want to do unpaid work on their own terms must qualify as end-stage capitalism...?
sounds like an excellent time to make an alternative
Maybe it's because I'm not a native English speaker but I can't really figure out if they are against or for AI answers?
They say ChatGPT is s parrot leading to poor quality, so I assume they are against, but Stack Overflow did indeed ban ChatGPT messages? Then they say AI detectors have many false positives, so I guess they are against strong filtering? So what are they for then? This piece needs a tl;dr or bullet list... So yeah I asked our large language friend:
Stances from the text:
A general moderation strike is being initiated.
The strike is in protest of recent and upcoming changes to policy and the platform by Stack Exchange, Inc.
Striking community members will refrain from moderating and curating content.
Critical community-driven anti-spam and quality control infrastructure will be shut down.
The new policy on AI-generated content is harmful and overrides community consensus.
There has been a serious failure to communicate on the part of Stack Exchange, Inc.
AI-generated content poses risks to the integrity of the platform and represents an honesty issue.
Stack Exchange, Inc. has ignored the needs and consensus of the community and made decisions without consulting those most affected.
The striking users want the AI policy change to be retracted or modified to address concerns and empower moderators.
They want the internal AI policy given to moderators to be revealed to the community.
Clear and open communication from Stack Exchange, Inc. regarding policy changes is demanded.
Collaboration with the community instead of fighting it is expected.
Stack Exchange, Inc. should be honest about the company's relationship with the community.
A change in leadership philosophy toward the community is needed.
Leadership should allocate resources based on community needs and involve the community in feature development.
Neglecting and mistreating volunteers can lead to a decrease in goodwill and motivation.
The concerns laid out in the open letter and the post should be addressed to end the strike.
Imho this is the key thing: They want the internal AI policy given to moderators to be revealed to the community, and in general want a more open/equal relationship with SO management.