Large language models reduce public knowledge sharing on online Q&A platforms
academic.oup.com
academic.oup.com
E.g. you can't just synthetically generate code, something or someone needs to run it and see if it performs the functions you actually asked of it.
You need to feed the LLM output into some kind of formal verification system, and only then add it back to the synthetic training dataset.
Here, for example - dumb recursive training causes model collapse:
I suspect most foundational models are now knowingly trained on at least some synthetic data.
I don't see how you can "bootstrap a smarter model" based on synthetic data from a previous-gen model this way. You may as well well just train your new model on the original training data.
The Information in the data isn't just about the output but its rate of occurrence/distribution. If what your base model has learnt is only enough to have the occasional flash of brilliance say 1 out of 40 responses and you are able to filter out these responses and generate as much as you like then you can very much 'bootstrap a better model' by training on these filtered results. You are only getting a function of the model's parameters if you train on its unfiltered, unaltered output.
Interesting - I actually think they perform quite well on code, considering that code has a set of correct answers (unlike most other tasks we use LLMs for on a daily basis). GitHub Copilot had a 30%+ acceptance rate (https://github.blog/news-insights/research/research-quantify...). How often does one accept the first answer that ChatGPT returns?
To answer your first question: new content is still being created in an LLM-assisted way, and a lot of it can be quite good. The rate of that happening is a lot lower than that of LLM-generated spam - this is the concerning part.
I agree, LLMs performance for coding tasks is super biased in favor of well-represented languages. I think this is what GitHub is trying to solve with custom private models for Copilot, but I expect that to be enterprise only.
Published documentation.
I'm going to make up a number but I'll defend it: 90% of the information content of stackoverflow is regurgitated from some manual somewhere. The problem is that the specific information you're looking for in the relevant documentation is often hard to find, and even when found is often hard to read. LLMs are fantastic at reading and understanding documentation.
From personal experience, I'm skeptical of the quantity and especially quality of published documentation available, the completeness of that documentation, the degree to which it both recognizes and covers all the relevant edge cases, etc. Even Apple, which used to be quite good at that kind of thing, has increasingly effectively referred developers to their WWDC videos. I'm also skeptical of the ability of the LLMs to ingest and properly synthesize that documentation - I'm willing to bet the answers from SO and Reddit are doing more heavy lifting on shaping the LLM's "answers" than you're hoping here.
There is nothing in my couple decades of programming or experience with LLMs that suggests to me that published documentation is going to be sufficient to let an LLM produce sufficient quality output without human synthesis somehwere in the loop.
I've answered dozens of questions on stackoverflow.com with tags like SIMD, SSE, AVX, NEON. Only a minority of these asked for a single SIMD instruction which does something specific. Usually people ask how to use the complete instruction set to accomplish something higher level.
Documentation alone doesn't answer questions like that, you need an expert who actually used that stuff.
For me, the current state of affairs is the difficulty of aiming search engines to the right answer. The answer might be out there, but figuring out HOW t oask the questions, which are the right keywords, etc - basically requires knowing where the answer is. LLMs have potential in rephrasing boththe question and what they might have glanced at here and there - even if obvious in hindsight.
Nothing will ever replace your own critical thinking and judgment.
> At least when discussing with real people after the arrows fly and the dust settles, we can figure out the truth.
You can actually do that with AI now. I have been able to correct AI many times via a Socratic approach (where I didn't know the correct answer, but I knew the answer the AI gave me was wrong).
yes, a lot of stuff is like that, and LLMs are a good replacement for searching the docs; but the most useful SO answers are regarding issues that are either not documented or poorly documented, and someone knows because they tried it and found what did or did not work
It’s helpful to realize the ways in which we do work the same way as AI, because it gives us perspective unto ourselves.
(I don’t follow regarding your true and false statement, and I don’t share your apparent pessimism about the fundamental limits of AI.)
Unlike freeform writing tasks, coding also has a strong feedback loop (i.e. does the code compile, run successfully, and output a result?), which means it is probably easier to generate synthetic training data for models.
I doubt it. Take a language like Rust or Haskell or even modern Java or Python. Without prolonged experience with the language, you have no idea how the various features interact in practice, what the best practices and typical pitfalls are, what common patterns and habits have been established by its practitioners, and so on. At best, the system would have to simulate building a number of nontrivial systems using the language in order to discover that knowledge, and in the end it would still be like someone locked in a room without knowledge of how the language is actually applied in the real world.
Cousin of the sufficiently smart compiler? :-p
Like "Help me to do this and that and use this list of internet resources to answer my questions"
Even limiting yourself to code generation, there are going to be a lot of software developers employed to write or generate code examples and documentation just for AIs to ingest.
I think eventually AIs will begin coding in programming languages that are designed for AI to understand and work with and not for people to understand.
The sheer difference in scale between the domain of “here are all the people in the world that have shared data publicly until now” and “here is the relatively tiny population of people being paid to add new information to an LLM” dooms the LLM to become outdated in an information hoarding society. So, the question in my mind is, “Why will people keep producing public information just for it to be devalued into LLMs?”
If you mean obfuscation, then yeah, maybe that makes sense to fit more into the window. But it’s easy to unobfuscate, usually.
Otherwise, I‘m not sure what the goal of an LLM specific language could be. Because I don’t feel most languages have been made purely to accommodate humans anyway, but they balance a lot of factors, like being true to the metal (like C) or functional purity (Haskell) or fault tolerance (Erlang). I‘m not sure what „being for LLMs“ could look like.
{/* rest of your functions here*} - DELETED
After a while it's only safe for doing tedious things like loops and switches.
So I guess our jobs are safe for a little while longer
It makes me wonder if the AI successfully found and digested obscure sources on the internet or was just better at making sense of the esoteric documentation than me. If the latter, perhaps the need for public samples will diminish.
But it takes so long and so much prompting skill to get to that point that the use cases seem limited.
We absolutely need public sources of truth at the very least until we can build systems that actually reason based on a combination of first principles and experience, and even then we need sources of truth for experience.
You simply cannot create solutions to new problems if your data gets too old to encompass the new subject matter. We have so systems which can adequately determine fact from fiction, and new human experiences will always need to be documented for machines to understand them.
and to be clear - as a layman, (in almost every field) I've recognized that llm's weren't up to the challenge of disavowing that notion from my phd friends up until o1 and never even tried, even though I've 'believed' since gpt 2
Any time I tried some of the more complex prompts I was working through something with Sonnet or 4o, o1 would totally miss important points and ignore a lot of the instructions while going really deep trying to figure out some relatively basic portions of the prompt.
Seems fine for reasonably simple prompts, but gets caught up when things get deeper.
The way I am connecting external content to the epub is through an anchoring system -- sequences of words can be hashed to form unique ids that are referenced by the enhancement. Doing this lets me index the enhancement content in such a way that is format-independent and doesn't require modifying the underlying epub.
The specific task o1 helped me with was determining what the text is visible at any given point in time. This text is then turned into hashes to pull the relevant enhancement content.
Getting the current words on the page doesn't seem all that complex, but the epub.js API is pretty confusing. There are abstractions like "Rendition", "Location", "Contents", "Range", and it's not always intuitive which of these will provide the appropriate methods. I'm sure I would have figured it out eventually by looking at the API docs, but GPT probably saved me and hour or two of trial-and-error.
My mom and grandma also read books with a physical notebook handy, and seems like modern technology should make that unnecessary.
I was laid off as part of the massive layoffs in 2022 and made a first pass at the project. I didn't use epubjs and instead wasted a lot of time making a custom rendering engine -- this was dumb. Eventually I got another job and paused the project. I started again in earnest a month or so ago using epubjs as the base.
hmm having a sort of mini-forum-like experience tied to particular pages in a book seems like a fascinating idea! being able to discuss plot twists and such only once you've already gotten to that point?
wow this seems like an amazing idea actually! any names yet? I'd love to check it out once it's done!
Obviously, not a given, but not unreasonable given what we have seen historically.
why would it be lucrative? The person paying would consider whether they'd get locked in to the framework/language, and be held hostage by the creator(s). This is friction to adoption. So LLMs will make popular, corporate backed languages/frameworks even more popular and drown out the small ones.
Scarcity of some knowledge. Not all knowledge exists on SO. You are right about the popular stuff, but the niche stuff will be like everything else niche, harder to get and thus more expensive. COBOL is typically used as an example of this, but COBOL was at least documented. I am saying this, because, while I completely buy that there will be executives who will attempt to do this, it won't be so easy to rewrite it all like Jassy from Amazon claims ( or more accurately, it will be easy, but with exciting new ways for ATMs, airlines and so on to mess with one's day).
<< The person paying would consider whether they'd get locked in
I want to believe a rational actor would do that. It makes sense. On the other hand, companies and people sign contracts that lock them in all the time to all sorts of things and for myriad of reasons including, but not limited to being wined and dined.
Again, I think you are right about the trend ( as it will exacerbate already existing issues ), but wrong about the end result.
Most of this "knowledge sharing on online Q&A platforms" is NOT creative activity. It's endless questions about the same issues everyone is having except the system developers themselves. Much of this is just displacing search platforms.
Seems like the humans using it can be also training and effectively paying for the right to do so? I would guess this could be easily gamified.
This underlying societal problem has knock-on effects just like chilling effects on free speech, but these effects we talk about are much more pernicious.
To understand, you have to know something about volunteer work, and what charity is. Charity is intended to help others who are doing poorly, at little to no personal benefit to you, to improve their circumstances. Its gods work, there is very little expectation involved on the giver. Its not charity if you are forced to give.
When you give something to someone as charity, and then that someone, or an observing third-party related to that charity extorts you and attempts to coerce you for more. What do you think the natural inclination is going to be?
There is now additional coercive personal cost attached to that which was not given freely; what naturally happens. This is theft.
Volunteer psychology says, you stop giving when it starts costing you more than you were willing to give. Those costs may not be monetary, in many cases they are personal, and many people who give often vet those who they are giving to so as to help good people rather than enable evil people. LLM's make no distinction.
Less people give, but that is not all. Less people who are intelligent give anything. They withdraw their support. Where the average person in society was once pulled up to higher forms of thought by those with intelligence and a gift for conveying thought, that vanishes as communication vanishes. From that point forward, the average is pulled only to baser and baser forms of thought and the average people seem like the smartest people in the room though that feeling is flawed (because they aren't that intelligent).
The exchange for charity is often I'll help you here, and you are obligated to use it to get yourself into a better position so you can help both yourself and other people along the way. Paying it forward if you can. Charity is wasted on those that would use what is given to support destructive and evil acts.
There are types of people who will just take until no more can be given which is a type of evil, which is why you need to vet people beforehand but an LLM is simply a programmatic implementation of taking. The entire idea of a robot worker replacing low-mid level jobs during limits of growth is one that seems enticing from the short-sighted producer side except it ends up stalling economic activity. Economies run through an equal exchange between two groups. The moment there is no exchange because money isn't being provided for labor, products can't be bought. Shortages occur, unrest, usually slavery then death when production systems for food break down.
We are already seeing this disruption in jobs in the Tech sector that is 5 times the national unemployment during peak hiring (where offpeak has hiring freezes), in a single year. When you can't keep talent, that talent has options and goes elsewhere, either contributing within the economy or creating their own (black markets). ECP in non-market socialist systems (like what we are about to have in about 5 years), guarantees a lot of death or slavery.
People volunteer to help others, if someone then uses those conversations to create a humunculus that steals that knowledge and gives it to anyone that asks, even potentially terrorist or other evil actors, no one who worked for the expertise will support that.
Inevitably it also means someone will use that same thing to replace those expert jobs entry level portions with a machine that can't reason, and no new candidates can get the experience to become an expert. Its a slow descent into an age of ruin that ends in collapse.
You end up with cyclical chaotic spirals of economic stagnation, followed by shortage, followed by slavery or death from starvation; no AGI needed. All that's needed is imposing arbitrary additional cost on those least able to bear it. Interference in the job market, making it impossible to find work, gender relations (making it impossible to couple up and have babies), political relations (that drive greater government and control at the same time limiting any future). Its dystopian but its becoming a reality because of evil people in positions that allow front-of-line blocking to course correct continue their evil work; and they are characteristically wllfully blind to the destructive acts they perform and the consequences.
There are a lot of low pay jobs that will just vanish, people don't have enough money to raise children as it is; old crowds out the new until finally, a great dying occurs. That is the future if we don't stop it and resolve the changes that were made (as the bridges were burnt) by the current generation in political power. They should have peaked in power in 95, exited in 2005-2010. They still hold the majority and its 2024 (30 years after taking power). They'll die of old age soon, and because they were the only ones that knew all the changes that were made, new bridges will have to be built from scratch with less knowledge, resources, goodwill/trust, and people. They claimed they wanted a better world for their children, while creating a hellscape for them instead made of lies and deceit.
Very few people are taught specifically what the social contract entails. Its well past time people learn, lest they end up being the supporting cause of their own destruction.
Thomas Paine's Rights of Man has a lot of good points on this.
Your first argument largely rests on this, but isn't the fact that so many LLM-created answers get posted actually proof of the opposite? The LLM is, after all, "giving" here - and it's not like those answers are useless, just because they're made by an LLM. LLMs are perfectly capable of giving good answers to a large number of common tech support questions, and are therefore ("voluntarily") helping people.
> We are already seeing this disruption in jobs in the Tech sector that is 5 times the national unemployment during peak hiring (where offpeak has hiring freezes), in a single year. [...] Inevitably it also means someone will use that same thing to replace those expert jobs entry level portions with a machine that can't reason, and no new candidates can get the experience to become an expert. Its a slow descent into an age of ruin that ends in collapse.
This is your other argument, which is slightly more solid - I think that at least in the short term, we are going to see many of those effects. But overall, this dystopian "collapse" scenario is economically naive, for at least two reasons:
1. As it becomes harder to enter professions requiring high degrees of specializiation as the entry-level positions dry up, the demand for real specialists in these fields increases accordingly. At some point, it becomes viable for new talent to enter those sectors even in the face of a low number of "true" entry-level positions, as there is now an incentive to accept a much steeper upfront cost in terms of time invested on learning the specialization. If there are jobs that can at all be done and that pay 500,000$ per year, somebody is going to do them.
2. And I mean, besides, the entire "ECP in non-market socialist systems [...] guarantees a lot of death or slavery" argument is plain wrong. ("Pure ideology! *sniffs*"). We already live in a post-market capitalist society. Amazon controls enough market share in a large number of sectors that their operating model resembles Soviet-style demand planning when calculating how much produce to order (that is, order to manufacture). The Economic Calculation Problem, in the age of computers, is solved. We are not playing with abacuses anymore.
At least the USSR's ideology claimed to represent the worker's interest, and so once that stopped happening in practice, the system collapsed. (And they were already in the age of computers, trying to do some of the first cybernetics.)
For the amount of responsibility to get things right, you now have jobs that have significantly less responsibilities or entry skills that are now competing with the same positions thanks to inflation. Who would bother signing up for that stress for no economic benefit when they can get any number of other jobs at 3/4 of the pay and none of the responsibility. This type of chaos is what you've discounted as not happening (ECP).
> The economic calculation problem ... is solved
Then why are we having chaotic disruptions, in growing frequency and amplitude, continually more concentrated business sectors based in Lenin's comments on attacking capitalism through debauching currency through inflation as he mentioned to Keynes (towards non-market socialism), and more importantly shortages that are starting to self-sustain at the individual level.
When common food goods are not getting replenished on a weekly basis, and they are 2-3x the normal price, there is an issue. When it hits 2-3 weeks it starts to compound as demand skyrockets for non-discretionary goods (and they can't keep up). Goods sell out almost immediately (as they have been more recently), and people start hoarding (which they are doing again with TP).
You don't seem to realize this is the problem in action. It is hardly solved, we aren't even to the non-market part yet and we are far enough along to see that chaos is increasing, distortions are increasing, and shortages are increasing and not being corrected within JIT shipping schedules.
Those three observations are cause for concern to stop and lookaround because chaos is a major characteristic of the ECP. Moving goods around doesn't perform economic calculation, and soviet-style demand planning failed historically once shortages occurred, the next things coming up in the near future is sustaining shortage and death from famine. India stopped exporting its food because they see this on the horizon and have a intimate history with famine.
The 1921 famine (bolshevik non-market) reduced national population by 4% YOY and dropped the birthrate negative; Mao killed off more than 33+ million during his non-market stint which caused famine. No accurate numbers to draw up a percentage YOY but it was significant.
Every single non-market socialist state has failed in the history of socialism, and that includes its derivatives. Its not appropriate to call it ideology when there is a historic record where not even a single non-market state exists today. The fact that they all had to have markets tied to capitalist states to get their prices shows the gravity of the issue, and now that the last capitalist state is falling to non-market socialism (through intentional design/sabotage), what do you suppose is going to happen to all those market-socialist countries that can no longer get their price data. Its already happening, and chaos is unpredictable; and inflation/deflation measures are limited in resolution (as a high hysteresis problem).
Non-market socialist states run into the slavery and death part, but you are not recognizing that, and discounting it because it hasn't happened to us yet (though its likely within 5 years for a point of no return), and it has happened to every non-market system.
The point of no return to non-market socialism is when debt growth exceeds GDP (outflows exceed inflows). This is structural of any ponzi scheme. Who decides, the producers. This is when rational producers abandon the market, and start shutting down because its no longer possible to make a profit in terms of purchasing power. From there its a short but inevitable jaunt to collapse.
Have you noticed that China's stimulus through printing isn't doing much? Every printing requires exponentially greater amounts, with diminishing returns.
I think you've sorely confused objective observation with ideology and discounted it without taking into account the factors rationally; these are factors showing a clear and present danger to all existing market-socialist states, as well as those capitalist bastions that will be pulled down in the near future from ongoing subversion.
Amazon doesn't produce, its a logistics company. Shortages self-sustain because producers are leaving the market, we saw this first with baby formula in 2020 when the last producer was closed, but what's coming will leave only state-run capacity for most sectors and industry.
Most knowledgeable people know these systems are brittle and parasitic and aren't capable of ECP for a myriad of already written and rational reasons by themselves.
Socialist production absent a market, doesn't work without slave labor. It doesn't even work longer-term either with slave labor because of inefficiencies at current population don't scale well with tech levels.
You might be fine claiming everything is fine now, despite a lot of evidence to the contrary that things are getting progressively worse, and will continue down that path because of these consequences, but what about in 5 years when you can't get food. Will it matter then, when you can't do anything about it, because the time to do something came and passed?
You think the Elites in China don't know this is going to happen? They are pushing so hard for the BRICS because they know this is going to happen, because they helped cause it in their dogmatic and blind fervor and orchestration.
Name one modern real non-market socialist state that exists today that has not failed, and I'll be willing to consider your last statement. As far as I know, none exist and that should frighten every single person today because that is where we are going whether we like it or not (as a function of system dynamics).
Your focus on the false dichotomy between state action and private action makes me want to go tongue-in-cheek and name France. Definitely a non-market socialist state.
And for the record, I do agree that everything will get progressively worse, for ideological and for material reasons. But in both cases for different reasons than you suggest. The unholy marriage of extreme concentrations of wealth and the reemergence of fascism has been blessed with many children, but the Musks, Murdochs, Kochs and Thiels of this world will have little recourse against an economic reality wrecked by climate change, the drying up of fossil fuels and the bitter demographic outlook.
Still though, you see high concentration, loss of jobs, high unemployment, all classic signs often seen in socialist states right before major calamity or outright failure.
I'd say though the France won't have gotten to non-market socialism until their primary underlying currency hits those milestones, but they will see the chaos much sooner when the shoe does drop.
There is plenty of data even without QA forums.
However their claim is weak and the effect is not quite as sensational as they make it sound.
First, they only present Figure 3 and not regression results for their suggested tests of LLMs being substitutes of bad quality posts. In contrast, they report tests for their random qualification by user experience (where someone is experienced if they posted 10 times). Now, why would they omit tests by post quality but show results by a random bucketing of user “experience”?
Second, their own Figure 3 “shows” a change in trends for good and neutral questions. Good questions were downtrending and now they are flat, and neutral questions (arguably the noise) went from an uptrend to flat. Bad question continue to go down, no visible change in the trend. This suggests the opposite, ie that LLMs are in fact substituting bad quality content.
I feel the conclusion needed a stronger statement and research doesn’t reward meticulous but unsurprising results. Hence the sensational title and the somewhat redacted results.
People are electing to not freely share information on public forums like they used to. They are retreating into discord and other services where they can put down motes and raise the draw bridges. And who can blame them? So many forums and social media sites and forums are engaging in increasingly hostile design and monetization processes, AI/LLM’s are crawling everywhere vacuuming up everything then putting them behind paywalls and ruining the original sources’ abilities to be found in search, algorithms designed to create engagement foster vitriol and controversy, the list goes on. HN is a rare exception these days.
So what happens? A bunch of people with niche interests or knowledge sets congregate into private communities and only talk to each other. Which makes it harder for new people to join. It’s a sad state of affairs if you ask me.
I don't know the first thing about this thing, help me get to where I know the first thing. This is not allowed any more.
I want to know the pros and cons of various things compared. this is not allowed.
I have quality questions regarding an approach that I know how to do, but I want to know better ways. This is generally not allowed but you might slip through if you ask it just right.
I pretty much know really well what I'm doing but having some difficulty finding the right documentation on some little thing,help me - this is allowe
Something does not work as per the documentation, help me, this is allowed
I think I have done everything right but it is not working, this is allowed and is generally a typo or something that you have put in the wrong order because you're tired.
At any rate, the ones that are not allowed are the only questions that are worth asking.
The last two that is allowed I generally find gets answered in the asking - I'm pretty good in the field I'm asking in, the rigor of making something match SO question requirements leads me to the answer.
If I ask one of the interesting disallowed questions and get shit on then I will probably go through a period of screw it, I will just look extra hard for the documentation before I bother with that site again.
Pro/Con questions are too likely to involve opinion and degenerate into flamewars. Some could be answered factually, but mostly are not. Others have no clear answers.
>Often these are students asking for someone to solve their homework problems.
I don't think I've been in any class since elementary school in which I did not have foundational knowledge, I'm talking "I just realized there must be a technical discipline that handles this issue and I can't google my way to it level of questions."
If I'm a student, I have a textbook and the ability to read. I'm not asking textbook readable or relevant literature readable in the thing I am studying questions because I, being in a class on the subject I would "know the first thing" to quote my earlier post, that first thing being how to get more good and relevant knowledge on the thing I am in a class in.
I'm talking about things you don't even know what questions to ask to get that foundational knowledge which is among the most interesting questions to ask - the problem with SO is it only wants me to ask questions in a field in which I am already fairly expert but I have just hit a temporary stumbling block for some reason.
I remember when I was working on a big government security project and there was a Java guy who was an expert in a field that I knew nothing about and he would laugh and say you can't go to SO and ask about how do I ... long bit of technical jargon outside my field that I sort of understood hung together, maybe eigenvectors came up (this was in 2013)
Second thing, yes I know SO does not want people to ask non-factual questions, and it does not want me to ask questions in fields in which I am uninformed, so it follows it wants me to ask questions that I can probably find out myself one way or another, only SO is more convenient.
I gave some reasons why I do not find SO particularly convenient or useful given their constraints implying this is probably the same for others, you said two of my reasons were no good, but I notice you did not have any input on the underlying question of - why are people not asking as many questions on SO as they once did?
I don't know why SO questions are declining -- perhaps people find SO frustrating, as you seem to, and they give up. I myself have never posted a question on SO as I generally have found that my questions had already been asked and answered. And lately, perhaps LLMs are providing better avenues for the sorts of questions you describe. That seems very plausible to me.
> If I'm a student, I have a textbook and the ability to read
You are such an outlier that I don't think you have the awareness to make any useful observations on this topic. Quite a lot of students in the US are now starting to lack the ability to read, horrifyingly (and it was never 100%), and using ChatGPT to do homework is common.
You went blind in one eye to see out the other.
FWIW, I found LLMs to be actually really good at those basic questions where I'm at expert at language X and I ask how to do similar thing in Y, using Y's terms (which might be named differently in X).
I believe this actually would work well:
- extra basic things, or things that depend on opinion etc: ask LLMs and let they infer and steer you
- advanced / off the beaten path questions that LLMs hallucinate on: ask on SO
Since there's no way to appeal duplicate close votes on SO until you have a pretty large amount of rep, this kinda creates a problem where there's a "silent mass" of duplicate questions that aren't really duplicates.
A basic example is this question: https://stackoverflow.com/q/27957454 , which is about disabling PRs on GitHub on the surface. The body text however reveals that the poster is instead asking how they can set up branch permissions and get certain accounts to bypass them.
I can already assure you that just by basic searching, this question will pop up first when you look up disabling PRs, and the accepted answer answers the question body (which means that it's almost certain a different question has been closed as a duplicate of this one), rather than the question title. You could give a more informative answer (which kinda happened here), but this is technically off-topic to the question being closed.
That's where SO gets it's bad rep for inaccurate duplicate closing from.
It's certainly not frustrating for me, I ask a question maybe once a year on SO, most of their content is, in my chosen disciplines, not technically interesting, it is no better than looking up code snippets in documentation (which most of the time is what it really, really is)
I suppose it's frustrating for SO that people no longer find it worthwhile to ask questions there.
>advanced / off the beaten path
show me an advanced and off the beaten path question that SO has answered well, that is just not worth the effort to try to get an answer - if you have an advanced and off the beaten path question that you can't answer then you ask it on SO just "in case" but really you will find the answer somewhere else or not at all in my experience.
This may have been allowed in like the first year while figuring out what kind of moderation worked, but it hasn't been as least since I started using it in like 2011. They just kept slipping through the cracks because so many questions are constantly being posted.
If you ask a question (i.e., add content to the wiki) that is not in scope, then of course it will get removed.
This doesn't make the slightest bit of sense. The people who would be concerned are the ones who are providing answers. They are not coming to SO solely to get answers.
Exact same experience here. Plus, being able to talk to maintainers directly has been great!
There’s some weird blind spot with techies who are unable to see the appeal. UX matters in a “the medium is the message”-kind of way. Also, GitHub is only marginally more open than discord. It’s indexable at the moment, yes, but would not surprise me at all if MS is gonna make an offensive move to protect “their” (read our) data from AI competitors.
I've been wondering the same thing recently. It's really inefficient for me to communicate with my fellow maintainers through Github discussions, issues and pull request conversations so my go-to has been private discord conversations. This is actually kind of inefficent since most open source repos will always have a bigger community on github vs on discord (not to mention that it's a hassle when some maintainers are Chinese and don't have access to Discord...)
Each platform fills different role, GitHub is project specific and access to maintainers, who do not monitor SO. Discord is chat, with maintainers and community. SO misses the connection to the people building and using the software, it's more like the Quora of software nowadays. That Discord and GitHub have terrible search is not detracting from their other benefits
2024: Discord is not indexed by AI slop generators, it's great
I reckon that, ultimately, the success of "gift culture", "kNoWlEdGe ShOuLd Be Free!!!1", F/OSS, et al. via the WWW will cast this entire Stallman-esque hacker ethos in a bad light.
We all work for IBM^H^H^HOpenAI, but there's nothing like the GPL to back us up now!
Large language models are used to aggregate and interpolate intellectual property.
This is performed with no acknowledgement of authorship or lineage, with no attribution or citation.
In effect, the intellectual property used to train such models becomes anonymous common property.
The social rewards (e.g., credit, respect) that often motivate open source work are undermined.
That's how it ends.
These LLM users produce the new training data. They are being assimilated into the tool.
This is the future of "open source": Anonymous common property continuously harvested from, and distributed to, LLM users.
The cost of contributions falls dramatically, eg, $100 is 200M tokens of GPT-3.5; so you’re talking enough to spend 10,000 tokens developing each line of a 20kloc project (amortized).
That’s a moderate project for a single donation and an afternoon of managing a workflow framework.
Open source as we know it today, not so much.
If LLMs will be the end of open source, then they will constitute that end for exactly the reason you write:
> Large language models are used to aggregate and interpolate intellectual property.
> This is performed with no acknowledgement of authorship or lineage, with no attribution or citation.
> In effect, the intellectual property used to train such models becomes anonymous common property.
And if those things are true and allowed to continue, then any IP relying on copyright is equally threatened. That could of course be the case, but it's hardly unique to open source. Open source is no different, here. Or are you suggesting that non-open-source copyrighted material (code or otherwise) is protected by keeping the "source" (or equivalent) secret? Good luck making money on that blockbuster movie if you don't dare show it to anyone, or that novel if you don't dare let people read it.
> The social rewards (e.g., credit, respect) that often motivate open source work are undermined.
First of all: Those aren't the only social rewards that motivate open source work. I'd even wager they aren't the most common motivators. Those rewards seem more like the image that actors that try to social-network-ify or gamify open source work want to paint.
Second: Why would those things go away? The artistic joy that drives a portrait painter didn't go away when the camera was invented. Sure, the pure monetary drive might suffer, but that drive is perhaps the drive that's least specific to open source work.
I think that is because, overall, the human nature does not change that much.
<< Open source is no different, here. Or are you suggesting that non-open-source copyrighted material (code or otherwise) is protected by keeping the "source" (or equivalent) secret? Good luck making money on that blockbuster movie if you don't dare show it to anyone, or that novel if you don't dare let people read it.
You may be conflating several different media types and we don't even know what the lawsuit tea leaves will tell us about that kind of visual/audio IP. As far as code goes, I think most companies have already shown how they protect themselves from 'open' source code.
I see this as a temporary problem however because LLMs are transitional. At some point it won't be necessary to train an LLM on the entirety of Reddit plus everything else ever written because there are obvious limits to statistical models like this and, as a counter point, that's not how humans learn. You may have read hundres of books in your life, maybe even thousands. You haven't read a million. You don't need to.
I find it interesting that this issue (which is theft, to be clear) is being framed as theft from the site or company that "owns" that data, rather than theft from the users who created it. All these user-generated content ("UGC") sites are doomed to eventually fail because their motivations diverge from their users and the endless quest to increase profits inevitably drives users away.
Another issue is how much IP consumption constitutes theft? If an LLM watches every movie ever made, that's probably theft. But how many is too many? Like Apocalypse Now was loosely based on or at least inspired by Heart of Darkness (the novel). Yet you can't accuse a human of "theft" by reading Heart of Darkness.
All art is derivative, as they say.
I agree but I think it may be privileging the human intelligence mechanism a bit too much. These LLMs are polymaths that can spit out content at a super human rate. It can generate poetry and literature similarly to code and answers about physics and car repair. It’s very rare for a human to be able to do that especially these days.
So I agree they’re transitional but only in the sense that our brains are transitional from the basal ganglia to the neocortex. In that sense I think LLMs will probably be a part of a future GAI brain with other things tracked on, but it’s not clear it will necessarily evolve to work like a human’s brain does.
Do you mean in theory or currently? Because currently, LLMs make simple errors (eg [1]) and are more capable of spitting out, well, nonsense. I think it's safe to say we're a long way from LLMs producing anything creatively good.
I'll put it this way: you won't be getting The Godfather from LLMs anytime soon but you can probably get an industrial film with generic music that tells you how to safely handle solvents, maybe.
Computers are generally good at doing math but LLMs generally aren't [2] and that really demonstrates the weaknesses in this statistical approach. ChatGPT (as one example) doesn't understand what numbers are or how to multiply them. It relies seeing similar answers to derive a likely answer so it often gets the first and large digits of the answer correct but not the middle. You can't keep scaling the input data to have it see every possible math question. That's just not practical.
Now multiplying two large numbers is a solvable problem. Counting Rs in strawberry is a solvable problem. But statistical LLMs are going to have a massive long tail of these problems. It's really going to take the next generational change to make progress.
[1]: https://www.inc.com/kit-eaton/how-many-rs-in-strawberry-this...
[2]: https://www.reachcapital.com/2024/07/16/why-llms-are-bad-at-...
I still think transformers and LLMs will likely remain as some component within that next gen architecture vs something completely alien.
I think that adding different forms of assistance remains the most interesting pattern right now.
https://chatgpt.com/share/670bfdbd-8624-8011-bc31-2ba66eab3e...
I didn't realize that it had come that far with the delegating of those problems to the code writing and executing part of itself.
> Another issue is how much IP consumption constitutes theft? If an LLM watches every movie ever made, that's probably theft.
It's hard to reconcile those two views, and I don't think theft is defined by "how much" is being stolen.
So I think future AI has to read also everything if it is used like ChatGPT these days by average people to ask for advice about almost anything.
Sometimes online forums are the only place where you can find solutions for niche situations and edge cases. Tricks which would have been very difficult to figure out on your own. LLMs can train on the official documentation of tools l/libraries but they can't experiment and figure out solutions to weird problems which are unfortunately very common in tech industry. If people stop sharing such solutions with others, it might become a big problem.
That's the most valuable aspect of it. When you find yourself in these niches situations, it's nice when you see someone has encountered it and has done the legwork to solve it, saving you hours and days. And that's why Wikis like the Arch Wiki are important. You need people to document the system, not just individual components.
LLMs train on way more than just the official documentation: they train on the code itself, the unit tests for that code (which, for well written projects, cover all sorts of undocumented edge-based) and - for popular projects - thousands of examples of that library being used (and unit tested) "in the wild".
This is why LLMs are so effective at helping figure out edge-cases for widely used libraries.
The best coding LLMs are also trained on additional custom examples written by humans who were paid to build proprietary training data for those LLMs.
I suspect they are increasingly trained on artificially created examples which have been validated (to a certain extent) through executing that code before adding it to the training data. That's a unique advantage for code - it's a lot harder to "verify" non-code generated prose since you can't execute that and see if you get an error.
If understanding the code was enough, we wouldn't have any bugs or counterintuitive behaviors.
> and - for popular projects - thousands of examples of that library being used (and unit tested) "in the wild".
If people stopped contributing to forums, we won't have any such data for new things that are being made.
I would argue that code in GitHub is much more useful, because it's presented in the context of a larger application and is also more likely to work.
'Listening to different spiritual gurus saying the same thing using different words' is like 'watching the same coloured glass pieces getting rearranged to form new patterns in kaleidoscope'
I've been thinking about this a lot lately. Could we train an AI, e.g. using RL and GAN, where it gets an IT task to perform based on a body of documentation, such that its fitness would then be measured based on both direct success on the task, and on the creation of new (distilled and better written) documentation that would allow an otherwise context-less copy of itself to do well on the task?
They're not visiting Stack Overflow for well-read material (the popular languages) because perplexity.ai, ChatGPT, Claude, etc., not only answer questions better than reading pages of StackOverflow, but will cut and paste the (right or wrong) answer for you faster.
If you're not on StackOverflow while asking questions, you aren't answering things there either. Nothing else is needed to explain their observations.
Of course, this means that to compete, StackOverflow and other Q&A forums must bring the answering usability (integrating answers into one's flow) to the top.
AI won't "answer questions better". It will cut out the middleman of parsing your question and matching it up with words in the shape of an answer. It will also frequently hallucinate, and essentially never sanity-check what you're trying to do.
But the main thing giving it a speed/convenience advantage over Q&A forums is that it doesn't give a damn about whether your question, and its answer, could help someone else (by being discoverable with a search engine and understood by that other person as being the same question, and being focused on a single issue). Aside from not being designed to care, it wouldn't benefit: it can just re-generate the same answer content for the next person who asks, in a different low-quality way. Unlike a human expert, AI will never tire of that task.
I am active in a few specialized fields and already I have to defined my advice against poorly crafted prompt responses.
I'm surprised when people don't already engage in questioning like that.
I've had to be doing it for decades at this point.
Much of the worst advice and information I've ever received has come from expensive human so-called "professionals" and "experts" like doctors, accountants, lawyers, financial advisors, professors, journalists, mechanics, and so on.
I now assume that anything such "experts" tell me is wrong, and too often that ends up being true.
Sourcing information and advice from a larger pool of online knowledge, even if the sources may be deemed "amateur" or "hobbyist" or "unreliable", has generally provided me with far better results and outcomes.
If an LLM is built upon a wide base of source information, I'm inclined to trust what it generates more than what a single human "professional" or "expert" says.
if i need advice on repairing a weird unique metal piece on a 1959 corvette, im going to trust the advice of an expert in classic corvettes way before i trust the advice of my barber who knows nothing about cars but confidently tells me to check the tire pressure.
this “oh no, experts have be wrong before” we see so much is wild to me. in nuanced fields i’ll take the advice of experts any day of the week waaaaaay before i take the advice from someone who’s entire knowledge of topic comes from a couple twitter post and a couple of youtube’s but their rhetoric sounds confident. confidently wrong dipshits and sophists are one of the plagues of the modern internet.
in complex nuanced subjects are experts wrong sometimes? absofuckinlutely. in complex nuanced subjects are they correct more often than random “did-my-own-research-for-20-minutes-but-got-distracted-because-i-can’t-focus-for-more-than-3-paragraphs-but-i-sound-confident guy?” absofuckinlutely.
does this mean you trust complete randoms just as much?
-----
Personally i trust the consensus, not necessarily one random person.
I have the same problem as the guy above, at this point i assume doctors are almost always wrong if the problem isn't something really common for specific enough
A bit like if you ask an LLM to tell you a joke they all tend to go with the same one
Which makes me think, maybe a good move for Stack Overflow (which does not allow the submission of LLM-generated answers, wisely imo) would be to add an AI agent that would suggest an answer for each question, that people could vote on. That way you can elicit human and machine answers, and still have the verification process.
- AI answers are grammatically correct and verbose so it looks like the poster put effort into it which deceives people into thinking the answers is more trustworthy than it is.
Barring trolls, humans (for the most part) only answer if they think they're right, and the more effort put into the answer the more likely they don't get it wrong.
Then also have agents that automatically give upvotes for used solutions. Weird world.
I’m just imagining the precogs talking to each other in Minority Report if that makes sense.
Best bet is to book publish, and require a license from anyone that wants to train on it.
Going back to the original source: By giving an answer to somebody on a Q&A site, they might be a kid learning and then building solutions I benefit from later, again. Similar with software.
And I also consider the total gain of knowledge for our society at large a gain.
While my marginal cost form many things is low. And often lower than a cost-benefit calculation.
And some Q&A questions strike a nerve and are interesting to me to answer (be it in thinking about the problem or in trying to boiling it down to a good answer), similar to open source. Some programming tasks as fun problems to solve, that's a gain, and then sharing the result cost me nothing.
Obeying the law is unethical. Think of the poor laid off police officers.
The fact that information can be duplicated with zero cost is a feature, not a bug.
Not sure why they shut down jobs; they recently brought back a poorer version of it.
For example, I was wondering if its okay to move my home directory to a different filesystem altogether and create a symlink from /home/. Where do I ask such questions? The freaking ZFS mailing list? SO? It was just a passerby question, and what I wanted more than the answer is the sense of community.
The only place that I know that have a wide enough range of interest, with many people that each know some of these stuff quite deep, is public, is easily accessible is unironically 4chan /g/.
I would rather go there then Discord where humanity's knowledge will be piped to /dev/null.
This would blogging 5.0. Or web 7.0.
Lo and behold, the answer is very polite, explanative and even lists common mistakes. It even added two very helpful user comments!
I asked it again how this answer would look in 2024 and it just updated the answer to the latest c++ standard!
Then! I asked it what a moderator would say when they chime in. Of course the moderator reminded everyone to stay on focus regarding the question, avoid opinions and back their answer by documentation or standards. In the end the mod thanked for everyone's contribution and keeping the discussion constructive!
Ah! What a wonderful world ChatGPT is living at! I want to be there too!
With ChatGPT I’m waiting seconds or even 10+ seconds for a response, it can be 100% wrong and I then have to readjust my question or point out how it’s wrong which takes me a while.
If the Q & A isn't our there, then it'll be asked on an online platform.
And don't say they wouldn't, we've already watched business decisions be made to pull back on spending to produce official documentation as the availability of the internet grew to the point vendors understood that everyone has always-online internet connections.
Half joking, but I am pretty tired of SO pedantry.
Q: how do I do thing X in C?
A: Why do you need to know this? The C standard doesn't say anything about X. The answer will depend on your compiler and platform. Are you sure you want to do X instead Y? What version of Ubuntu are you running?
My message contained greetings and the question in the same message. I was banned immediately and the response from the mods was:
- do not greet; we don't have time for that bullshit
- do not use natural language questions; submit a test case and we will understand what you mean through your code
- do not abbreviate words (you have abbreviated "you" as "u"); if you do not have time to type the words, we do not have time to read them
The ban lasted for a week! :D
Don't even get me started about how dumb rule 2 is, though. And rule 3 doesn't even work for normal English as many things are abbreviated, e.g. this example.
And of course, you didn't greet and wait, you just put a pleasantry in the same message. Jeez.
I'm 100% sure I'd never have gone back after that rude ban.
upon saying this, the young apprentice was enlightened
I'm pretty sure that "rule" was more aimed towards "just ask your question" rather than "greet, make smalltalk, then ask your question".
I have similar rules, though I don't communicate them as aggressively, and don't ban people for breaking them, I just don't reply to greetings coming from people I know aren't looking to talk to me to ask me how I've been. It's a lot easier if you send the question you have instead of sending "Hi, how are you?" and then wait for 3 minutes to type out your question.
Yep, it's in there with "please use the public channel rather than DMs".
Some nettiquite, please.
Of course, in the process of typing out their question, they usually find the answer on their own.
That is really absurd! AFAIK, it is not possible to pose a question to a human in C++.
This level of dogmatism and ignorance of human communication reminds me of a TL I worked with once who believed that their project's C codebase was "self-documenting". They would categorically reject PRs that contained comments, even "why" comments that were legitimately informative. It was a very frustrating experience, but at least I have some anecdotes now that are funny in retrospect.
I will no longer work anywhere that has this kind of culture.
Here's a made-up example:
// Parses a Foo message [spec]. Returns true iff the parse was successful.
// The `out` and `in` pointers must be non-null.
//
// [spec]: https://foo.example/spec
bool parse_foo(foo_t* out, const char* in, size_t in_len) {
assert(out);
assert(in);
// TODO(tracker.example/issue#123) Delete the Foo v1 parser. Nobody will
// be sending us Foo v1 message when all the clients are upgraded to Bar
// v7.1.
//
// Starting in Foo v2, messages begin with a two-byte version identifier.
// Prior to v2, there was no version tag at all, so we have to make an
// educated guess.
//
// For more context on the Foo project's decision to add a version tag:
// https://foo.example/specv2#breaking-change-version-tag
uint16_t tag_or_len;
if (!consume_u16(&tag_or_len, &in, &in_len)) {
return false;
}
if (tag_or_len == in_len) {
return parse_foo_v1(out, in, in_len);
}
return parse_foo_v2(out, in, in_len);
}
On that old project, I believe the TL would have rejected each of these comments on the grounds that the code is self-documenting. I find this to be absurd:(1) Function comments are necessary to define a contract with the caller. Without it, callers are just guessing at proper use, which is particularly dangerous in C. Imagine if a caller guessed that the parser would gracefully degrade into a validator when `out` is NULL. (This is a mild example! I'm sure I could come up with an example where the consequence is UB or `rm -rf /`.)
(2) The TODO comment links to the issue tracker. There's no universe in which a reader could have found issue#123 purely by reading the code. On the aforementioned project, we were constantly rediscovering issues after wasting hours/days retreading old territory.
(3) Regarding the "Starting in Foo v2" comment... OK, fine, maybe this could have been inferred from reading the code. But it eases the reader into what's about to happen while providing further context with a link. On balance, I think this kind of "what" comment is worth including even though it mirrors the code.
> - do not greet; we don't have time for that bullshit
and
> do not abbreviate words (you have abbreviated "you" as "u"); if you do not have time to type the words, we do not have time to read them
So they have apparently enough time to read full words, it seems!
edit: I remember how some communities changed into: The help isn't good enough, you should help harder, I want you to help me by these conventions. Then they leave after getting their answer and no one has seen them ever again rather than join the help desk.
With “u”, I have to pause for a moment and think “that’s not a normal word; I wonder if they meant to type ‘i’ instead (and just hit a little left of target)?” and then maybe read the passage twice to see which is more likely.
I don’t think it’s quite as much a contradiction. (It still could be more gruff than needed.)
Now that Rust is succesfully assimilating those communities, I have noticed the same toxicity on less well moderated forums, like the subreddit. The Discord luckily is still great.
It's probably really important to separate the curmudgeons from the fresh initiates to provide an enjoyable and positive experience for both groups. Discord makes that really easy.
In the Ruby IRC channel curmudgeons would simply be shot down instantly with MINASWAN style arguments. In the Haskell IRC channel I guess it was basically accepted that everyone was learning new things all the time, and there was always someone willing to teach at the level you were trying to learn.
If you're out there-- how does it feel to know that what you meant as a efficient course-correction for newcomers was instead a social shaming that cut so deep that the message you wrote is still burned verbatim into their memory after all these years?
To be clear, I'm taking OP's experience as a common case of IRC newbies at that time on many channels. I certainly experienced something like it (though I can't remember the exact text), and I've read many others post on HN about the same behavior from the IRC days.
Edit: clarifications
Maybe that was the point?
Oh my, this reminded me how some 20 years ago I was a high school kid and dared to install a more nerdy Linux distro (which I won't name here) on my home computer. After some big upgrade, the system stopped booting, and when in panic I asked for help at the official forum, I got responses that were shaming me for blindly copying commands from their official website without consulting some README files. That's how I switched to Debian and never looked back.
I've never been a IRC mod for a channel of anywhere near that significance. But as someone who has put in considerable effort for the last few years trying to explain basic fundamental ideas about what Stack Overflow is and how it's supposed to work to people (including to people whose accounts are 15+ years old but who insist on continuing to treat the Q&A like a discussion forum)...
... I'm kinda envious.
Being presented with a text entry box, and the implied contract that what you type there will be broadcast to a wide audience, is not supposed to imply a right to reject the community's telos and substitute your own. Gatekeeping of this sort is important; otherwise you end up with the Wikipedia page for "Dog" being flooded with people trying to get free veterinary consultation. Communities are allowed to have goals and purposes that aren't obvious and which don't match the apparent design, and certainly they aren't required to have goals which encompass everything possible with the sites software.
There are countless places on the Internet that work like a discussion forum where you can post a question, or a request for help, or just a general state of confusion, about a problem where you're just completely lost; and where you can expect to have a back-and-forth multi-way conversation with others to try and diagnose things or come to a state of understanding; and where nobody cares if your real concern ever surfaces in the form of an explicit, coherent question; and where there's no expectation that any of this dialogue should ever be useful to anyone else.
Stack Overflow is not that, explicitly and by design; and that came about specifically so that people who do have some clue and want to find an answer to an actual question, can do so without having to pick through a dialogue of the sort described above, following an arbitrarily long chain of posts of questionable relevance (that might well end in "never mind, I fixed it" with no explanation).
But in order to become a repository of such questions, people presented with the question submission form... need to be restricted to asking such questions.
Sometimes, I really want to do X, I know it may be questionable, I know the safest is "probably don't want to do it", and yet, that's not someone else's (or LLMs) business, I know exactly what I want to do, and I'm asking if anyone knows HOW, not IF.
So IMO it's not a flaw, it's a very useful feature, and I really do hope LLMs stay that way.
That’s when I stopped investing any effort into that community.
Turned out it, counter-intuitively, was impossible. And not documented anywhere.
After some point it came to a point that if you're not asking a complete problem which can't be modeled as a logic statement, you're labeled as stupid for not knowing better. The thing is, if I knew better or already found the answer, I'd not be asking to SO in the first place.
After a couple of incidents, I left the place for the better. I can do my own research, and share my knowledge elsewhere.
Now they're training their and others models with that corpus, I'll never add a single dot to their dataset.
Someone who tells you something along the lines of "this is is common misconception, maybe you're looking for X instead" is being helpful and kind, and is not at all "labeling you as stupid". On Stack Overflow, it's not at all required that you actually need the question answered, as asked, in order to ask it. (You're even actively encouraged to ask and answer your own questions, as long as both question and answer meet the usual standards, as a way to share your expertise.)
It has never been acceptable to insult others openly on the site - but moderation resources are, and always have been, extremely strained, and the community simply doesn't see certain things as insulting or "mean" that others might. In particular, downvotes aren't withheld out of sympathy, because they are understood to be purely for content rating - but many users take them personally anyway. Curators commonly get yelled at when they comment (thus identifying themselves) to explain policy in polite but blunt copy-paste terms; and they get seethed at when they don't comment. There is a standard set of reasons to close questions (https://meta.stackoverflow.com/questions/417476), and a how-to-ask guide (https://stackoverflow.com/help/how-to-ask), and a site tour (https://stackoverflow.com/tour), and an entire Meta site with a variety of FAQ entries and other common references (for example, I wrote the "proposed faq" of https://meta.stackoverflow.com/questions/429808); but very few people seem interested in checking the extant documentation for the community norms, and then feel slighted when the community norms aren't what they assume they ought to be (because it works that way everywhere else). Communities should be allowed to have their own norms, so that they can pursue their own goals (https://meta.stackoverflow.com/questions/254770).
A ton of users confuse "know better or find the answer ahead of time" for debugging. Stack Overflow allows for what are commonly called "debugging questions", but this does not mean questions wherein you ask someone to debug the code for you. Why? Because that can't ever be helpful to someone else - nobody else has your code, and thus they can't benefit from someone locating the bug in your code. The "debugging questions" that Stack Overflow does want are questions about behaviour that isn't understood, after you have isolated the misbehaving part. These are ideally presented in the form of a "minimal reproducible example" or MRE (https://stackoverflow.com/help/minimal-reproducible-example).
Stack Overflow explicitly expects you to do research before asking your question (https://meta.stackoverflow.com/questions/261592) - because most of the goal of the site is to cover questions that can't be answered that way. Traditional reference documentation is only part of the puzzle (https://diataxis.fr/); the Q&A site format allows for explanations (for "debugging questions") and how-to guides (in response to simple queries about how to do some atomic, well-specified task). (Tutorials don't fit in this format because they don't start with a clear question from the student.)
SO does suck, but i've found that if you clarify in the question what you want, and pre-empt the Y instead of X type answers, you will get some results.
In other words, perhaps in your very specific case, your question is not XY problem but for the vast majority of visitors from google it won't be so. https://en.wikipedia.org/wiki/XY_problem
Personally, I always answered SO from at least two perspectives: how the question looks for someone coming from google and how the author might interpret it.
Increasingly, the free products of experts are stolen from them with the pretext that "users need to be protected". Entire open source projects are stolen by corporations and the experts are removed using the CoC wedge.
Now SO answers are stolen because the experts are not trained like hotel receptionists (while being short of time and unpaid).
I'm sure that the corporations who steal are very polite and CoC compliant, and when they fire all developers once an AGI is developed, the firing notices will be in business speak, polite, express regret and wish you all the best in your future endeavors!
Additionally, these platforms tend to attract a fair amount of spam (self promotion etc) which can make it very hard to find high-quality responses.
Then equipped with the right term it's way easier to find reliable information about what you need.
Maybe we can get this multi-headed advantage back from LLMs by applying a team of divergent AIs to the same problem. I've had other occasions when OpenAI gave me crap that Claude corrected, and visa versa.
- do a task
- criticize your job on that task
- redo that task based on criticism
I find giving the LLM a process greatly improves the results.
Although if I’m truly honest with myself, even after many years of developing, the true cycle of me writing code is: over confidence, then shock it didn’t work 100% the first time, wondering if there is a bug in the compiler, and then reality setting in that of course the compiler is fine and I just made my 15th off-by-one error of the day :)
I too can write made up criticism if that’s what my boss wants in the workplace — but that doesn’t suddenly invalidate my ability to criticize my own work to improve it.
I've been arguing with Copilot back and forth where it gave me a half-working solution that seemed overly complicated but since I was new to the tech used, I couldn't say what exactly was wrong. After a couple of hours, I googled the background and trust my instinct and was able to simplify the code.
At that situation, where I iteratively improved the solution by telling Copilot things seem to complicated and this or that isn't working. That led the LLM to actually come back with better ideas. I kept asking myself why something like you propose isn't baked into the system.
I don’t want to be that guy saying this, but 99% of the top results on google from Medium related to anything technical is literally the reworded/reframed version of the official quick start guide.
There are some very rare gems, but it is hard to find those among the above mentioned ocean of reworded quick starts disguised as “how to X”, “fixing Y”. Almost reminds me of the SEO junks when you search “how to restart iPhone” and find answers that dance around letting it die from battery drain and then charge, install this software, take it to the apple repair shop, go to settings and traverse many steps while not saying that if you are between these models use the power+volume up button trick.
End of rant.
Blogging platforms are the worst though. Medium looked pretty OK when it first came out. But now it is just a platform for self-promotion. Substack is like 75% of the way through that transition IMO.
People who do interesting things spend most of their time doing the thing. So, non-practicing bloggers and other influencers will naturally overwhelm the people who actually have anything interesting to report.
The best Q&A platform would be the one where experts and scientists answer questions but sites like Wikipedia and Reddit showed that broad range of audience can also be pretty good at providing useful information and moderating it.
But to me the worst issue is it's now "Dead Overflow": most answers are completely, totally and utterly outdated. And seen that they made the mistake of having the concept of an "accepted answer" (which should never have existed), it only makes the issue worse.
If it's a question about things that don't change often, like algorithms, then it's OK. But for anything "tech", technical rot is a very real thing.
To me SO has both outdated and inaccurate answers.
On the one hand, good new answers can remain buried under mediocre and outdated answers more or less indefinitely, because they received hundreds of questionable upvotes in the past. (Despite popular perception, the average culture of Stack Overflow is very heavily weighted towards upvoting. I have statistical analysis of this on the meta site somewhere.)
On the other hand, there are tons of popular reference questions that had to be "protected" so that people couldn't just sign up and contribute the 100th answer (yes, the count really does range into triple digits in several cases, especially if you have access to deleted answers) to a question where there are maybe five actually reasonable-considered-distinct answers possible. And this protection is a quite low barrier - 10 reputation, IIRC. (I'm not sure what motivates people to write these new answers. It might have something to do with being able to point to a Stack Overflow profile which boasts that you've "reached" millions of people, due to its laughably naive algorithm for that.)
Whatever they’re doing, it isn’t working.
Blame mods. Blame AI. Blame askers… whatever man.
That is a sinking ship.
If you don’t see people complain about SO, it’s because they aren’t using it, not because they’re using the search.
Pretty hard to argue at this point that the problem is with the users being too shit to use the platform.
That’s some high level BS.
Why all the hate? AI answers can suck, definitely. Stackoverflow literally holds the modern industry up. Next time you have a technical problem or error you don't understand go ahead and avoid the easy answers given on the platform and see how much better the web is without it. I don't understand, what kind questions do you have?
As for the AI, it can only erode the quality of the content on SO.
I had one question that was a bit odd and went against testing dogma that I had a friend post. He pulled it 30 minutes later as he was already down 30 votes. It was a thing that's not best practice in most cases but also in certain situations the only way to do it. Like when you're testing apis you don't control.
In some sections people also want textbook or better quality answers from random strangers on the internet.
The final part is that you at least used to have to build up a lot of karma to be able to post effectively or at all in some sections or be seen. Which is very catch 22.
So it can be both very useful and very sh*t.
- it's less likely that the question even gets shown
- it's less likely that people will even click on it
- it's less likely that people who think it's bad will bother to vote on it, since the votes are already doing the right thing
- if it's really bad, it will be marked for deletion before it gets that many downvotes anyway
SO has its problems but I don't even recognize half the things people complain about.
SO is not a pure Q&A site. It is essentially a wiki where the contents are formatted as Q&As, and asking questions is merely a method to contribute toward this wiki. This is why, e.g., duplicates are aggressively culled.
The thing is, the meta community of Stack Overflow - and of other similar sites like Codidact - generally understand "Q&A site" to mean the exact thing you describe.
The thing where you "ask a question"[0] and start of a chain of responses which ideally leads to you sorting out your problem, is what we call a discussion forum. The Q&A format is about so much more than labelling one post as a "question" and everything else as an "answer" or a "comment" and then organizing the "answers" in a certain way on the page.
[0]: which doesn't necessarily have a question mark or a question word in it, apparently, and which rambles without a clear point of focus, and which might be more of a generic request for help - see https://meta.stackoverflow.com/questions/284236.
It's working just fine. The decline in the rate of new questions is generally seen as a good thing, as it's a sign of reaching a point where the low-hanging fruit has been properly picked and dealt with and new questions are only concerned with things that actually need to be updated because the surrounding world has changed (i.e., new versions of APIs and libraries).
>Pretty hard to argue at this point that the problem is with the users being too shit to use the platform.
On the contrary. Almost everyone who comes to the site seems to want to use it in a way that is fundamentally at odds with the site's intended purpose. The goal is to build a searchable repository of clear, focused, canonicalized questions - such that you can find them with a search engine, recognize that you've found the right question, understand the scope of the question, and directly see the best answers to the highest-quality version of that question. When people see a question submission form and treat it as they would the post submission form on a discussion forum, that actively pollutes said repository. It takes time away from subject-matter experts; it makes it harder for curators to find duplicates and identify the best versions thereof to canonicalize; and it increases the probability that the next person with the same problem, armed with a search engine, will find a dud.
Doesn’t seem to be a great strategy to always need these things retrained, but OpenAI’s o1 has things from early 2024
Don’t ask about knowledge cutoffs anymore, that’s not how these things are trained these days. They don’t know their names or the date.
Not to mention there will be no data for the model to learn the new stuff anyway, since places like SO will get zero responses with the new stuff for the model to crawl
It is really difficult if you need project flags and configurations to make things work, instead of just code
Github issues gets crawled, where many of these frameworks have their community
Some of them are so overburdened that navigating all of the rules and expectations becomes a skill in itself. A single innocent misstep turns simple questions into lectures about how you’ve violated the rules.
One Slack I joined has created a Slackbot to enforce these rules. It became a game in itself for people to add new rules to the bot. Now it triggers on a large dictionary of problematic words such as “blind” (potentially offensive to people with vision impairments. Don’t bother discussing poker.). It gives a stern warning if anyone accidentally says “crazy” (offensive to those with mental health problems) or “you guys” (how dare you be so sexist).
They even created a rule that you have to make sure someone wants advice about a situation before offering it, because a group of people decided it was too presumptuous and potentially sexist (I don’t know how) for people to give advice when the other person may have only wanted to vent. This creates the weirdest situations where someone posts a question in channels named “Help and advice” and then lurkers wait to jump on anyone who offers advice if the question wasn’t explicitly phrased in a way that unequivocally requested advice.
It’s all so very tiresome to navigate. Some people appear to thrive in this environment where there are rules for everything. People who memorize and enforce all of the rules on others get to operate a tiny little power trip while opening an opportunity to lecture internet strangers all day.
It’s honestly refreshing to go from that to asking an LLM that you know isn’t going to turn your question into a lecture on social issues because you used a secretly problematic word or broke rule #73 on the ever growing list of community rules.
Toddlers go through this sometimes around ages 2 or 3. They discover the "rules" for the first time and delight in brandishing them.
The fundamental problem is that communities/forums (in the general sense, e.g., market squares) don't scale, period. Because moderation and (transmission and error correction of) social mores don't scale.
Maybe initially, but in the community I’m talking about rules are introduced to prevent situations that might offend someone. For example, the rule warning against using the word “blind” was introduced by someone who thought it was a good thing to do in case a person with vision issues maybe got offended by it at some point in the future.
It’s a small group of people introducing the rules. Introducing a new rule brings a lot of celebration for the person’s thoughtfulness and earns a lot of praise and thanks for making the community safer. It’s turned into a meta-game in itself, much like how I feel when I navigate Stack Overflow
Sounds like that one is on its way out anyway...
That said, when that is not available, LLMs do an excellent job or rubber ducky-ing complicated topics.
There have been two places that I remember where arrogance of the esoterati drive two feedback cycles:
1. People leave after seeking help for an issue they believed needed the input of masters.
2. Because of gruff treatment, the masters receive complaints and indignation, triggering a backfire effect feedback loop, often under the guise of said masters not wanting to be overwhelmed by common problems and issues.
There is a few practical things that can help with this (clear guides to point to, etc.), but the missing element is kindness / non-judgmental responsiveness.
I am Stack overflow mod, dealing with other mods. There is definitely unnecessary hostility there, but most of question closes and downvotes Go 90% to low quality questiond which lack proper professionalism to warrant anyone's time. It is remaining 10% that turns off people.
We can also take analogs from the death of Usenet.
Without seeing faces, people just don't do this very well.
Q&A was designed to handle the social explosion problem and the eternal September problems by having a larger percent of the username take an interest in the community over time and continue to maintain that ideal. Things like comments and discussions being difficult is part of the design to make it so that you don't get protracted discussions that in turn needs more moderation resources.
The fraction of the people doing the curation and moderation of the site overall has dropped. The reasons for that drop are manyfold. I believe that much of it falls squarely upon Stack Overflow corporate without considering second order effects of engaging and managing the community of people who are interested in the success of the site as they envision.
Ultimately, Stack Overflow has become too successful and the people looking to it now have a different vision for what it should be that comes into conflict with both the design of the site and the vision of the core group.
While Stack Overflow can thrive with a smaller number of people asking "good" (yes, very subjective) questions it has difficulty when it strays into questions that need discussion (which its design comes into conflict with) or too many questions for the committed core group to maintain. Smaller sites can (and do) have a larger fraction of the user base committed to the goals of the site and in turn are able to provide more individual guidance - while Stack Overflow has long gone past that point.
---
Stack Overflow and its Q&A format that has been often copied works for certain sized user bases. It needs enough people to keep it interesting, but it fails to scale when too many people participate who have a different idea of what questions should be there.
There is a lack of moderation tools for the core user base to be able to manage it at scale (you will note the history of Stack Overflow has been removing and restricting moderation tools until it gets "too" bad - see also removal of 20k users helping with flag handling and the continued rescoping of close reasons).
Until someone comes up with a fundamentally different approach that is able to handle moderation at scale or sufficient barriers for new accounts (to handle the Eternal September problem), we are going to continue to see Stack Overflow clones spout and die on the vine along with a continued balkanization of knowledge in smaller areas that are able handle vision and moderation at a smaller scale.
---
Every attempt at a site I've seen since (and I include things like Lemmy in this which did a "copy reddit" and then worry (or not) about moderation) have started from a "get popular, then work on the moderation problem" which is ultimately too late to really solve the problem. The tools for moderation need to be baked into the design from the start.
While they are certainly not perfect, they willingly spend their own spare time to help other peoples for free. I disagree with calling them arseholes.
The problems are two-fold:
1. Any community with volunteer moderators attracts the kind of people you don't want to be moderators. They enjoy rigidly enforcing the rules even if it makes no sense.
2. There are two ways to find questions and answer them: random new questions from the review queue, and from Google when you're searching for a problem you have. SO encourages the former, and unfortunately the vast majority of questions are awful. If you go and review questions like this you will go "downvote close, downvote close, downvote close". You're going to correctly close a load of trash questions that nobody cares about and a load of good questions you just don't understand.
I've started recording a list of questions I've asked that get idiotic downvotes or closed, so I can write a proper rant about it with concrete examples. Otherwise you get people dismissing the problem as imaginary.
These mods now hold SO hostage. SO is definitely aware of the problem but they can't instigate proper changes to fix it because the mods like this situation and they revolt if SO tries to remedy it.
Sure the mods were arseholes etc.. but before gpt never minded using it .
Except of course that reading Stackoverflow directly has better retention rates, better explanations and more in-depth discussions.
(My view is that moderators can be annoying but the issue is overblown.)
What I find is that the LLMs are not spitting out SO text word for word, as one would when plagiarizing. Rather, the LLM uses the context and words of my question when answering, making the response specific and cohesive (by piecing together answers from across questions).
It's a core issue in academia and other areas where the output is heavily the written word.
[0] https://en.wikipedia.org/wiki/Plagiarism#Self-plagiarism
[1] https://ori.hhs.gov/self-plagiarism
[2] https://www.aje.com/arc/self-plagiarism-how-to-define-it-and...
(1) If you are reading your entrypoint into an area of research, or a group, it is useful context on first encounter
(2) If you are not, then you can easily skip it
(3) Citing, instead of reproducing sections like background work, means you have to go look up other papers, meaning a paper can no longer stand on its own.
Self-plagiarism is an opinion among a subset of academics, not something widely discussed or debated. Are there bad apples, sure. Is there a systemic issue, I don't think so.
No, you don't. You license it. The community gets a Creative Commons license (https://stackoverflow.com/help/licensing), and the company gets a few additional rights (https://stackoverflow.com/legal/terms-of-service/public#lice...). You retain copyright.
When that stopped, the site (and to some degree its growing number of sister sites) continued on its ballistic curve, slowly but continuously descending into the abyss.
My primary take away is that we have not found a way to scale moderation. SO was doomed anyway, LLMs have just sped up that process.
Compared to anything before it (endless phpBB forums, expertsexchange, etc.) it was just light years ahead.
Even today compared the SO UI with Quora. It's still 10x better.