The Fall of Stack Overflow
observablehq.com
observablehq.com
However, there is no commensurate decrease in posts/votes during the same time period. Posts/votes remained relatively constant through 2022 (modulo normal seasonal fluctuations), until February 2023 when both fell off a cliff (I assume due to the rise of LLMs). Traffic data are sourced from Google Analytics, while post/vote data are computed internally by StackOverflow [1]. I wonder if the apparent precipitous drop in traffic in May 2022 is simply an artifact of Google Analytics suddenly changing how it tracks traffic/visitors.
the post doesn't explain where they got these traffic numbers, and it seems unlikely they have access to stackoverflow's real traffic stats. they're using some sort of estimation here. there's always a chance that their estimates are wrong - especially if they're showing implausible shifts like this.
The question and user profile view stats (among other things) are public on the Data Explorer and included in the dumps:
https://meta.stackexchange.com/questions/2677/database-schem...
Not sure what they’re counting as a “visit”, though.
The new content rate been dropping at a dismally constant rate for a long time, but the first few months in 2023 were awfully grim. I wonder what might've corresponded to that.
The fall in the beginning of 2023 may be the introduction of ChatGPT. A more worrying idea is that the numbers reflect not just the decline of SO but a decline of the whole IT business.
I personally used to like StackOverflow as my last recourse: I grew up in those years where we had to RTFM, and I kept the habit. So if I go ask on StackOverflow, it is a tricky question. It used to be fine, and I was getting an answer eventually (sometimes after adding a bounty).
But in the last few years, I have had legit questions downvoted or even closed, and it was obvious that the people voting to close it did not even understand it. I agree that the moderation culture on StackOverflow is toxic. If everytime I contribute something, I have to fight to not get downvoted or closed, then I will slowly stop contributing.
"@JourneymanGeekOnStrike Yeah, if you go back further, the "traffic" numbers see a 40M/week drop between April and May 2022, which is when the cookie tracking changed, and then normalizes again until December. So, prior to the cookie changes, traffic was about 140-150M per week. But, to be clear - this is stuff we're aware of and have "corrected" for, I guess."
Nevertheless post the question and provide an answer. Everybody wins: you reap the upvotes, and everyone else benefits from the shared knowledge.
Unfortunately, there is a moderator strike now so I can't tell you to flag such an answer...
For 12 years, they have not figured this one out. New users will ask a very valid question and then won't respond anymore. I have seen this one played out every single day. Back in the day, users were generous with the upvotes even for a simple basic question, this is not the case anymore today.
Also SO is participation hostile unless you’re a pro, so as a newbie I’m not going to do anything other than ask and lurk, because I’m not worthy
On SO this hostility is pronounced because participants believe that if they get a lot of points they have easier time finding a well-paying job.
It’s true, though. I have a high-rep account and I’ve had a few jobs fall in my lap because of that.
It's just been a while since anyone has started trying to integrate tools with each other, outside of the established players.
reputation is meaningless and bloated in stackoverflow now there are many 100k reputation people because asking or answering basic shits on javascript/python/pandas/git
Then people will still try to close my question because it's a duplicate of B.
The issue is that all users are moderators and the newcomers are the ones not reading questions etc.
I would flag this sort of thing for a real moderator.
Then you point out that they completely disregarded what you wrote and blame you for misdirecting them. Then you get the downvotes.
The SO way.
especially YouTube links. sadly, it would not surprise me if these people are earning a decent enough money from ads to make it worth their while to be "content creators" solely from search results from Googs
If three segments of the internet think the same piece of information is relevant, that should affect the score of all 3 copies, not just the largest segment.
When content republished on some bizarre/sketchy/unaccountable ~adfarm outranks the site where it first appeared, users of Google's search service end up at higher risk of getting phished or infected with malware.
Is there some benefit you see here that outweighs this downside risk?
A few years ago, my programming-related queries would hit Stack Overflow as the first or second result. Now it's very frequently spammy garbage in the top 2-3 slots.
Remember, when people were hired because they knew the secret sauce on how to get the best Google ranking. Google experts?
Well, it turns out that the person at Google that was responsible for keeping the algorithm fresh and the search results fresh retired and everything went to shit when they left.
Actually, I’m betting that person did leave the company, but the real damage happened when someone came along and convinced everyone they knew the real trick to better search results and we have the shit that is now Google. Nice work new guy! Let me rephrase that. Nice work to the guy that thinks they are smarter than everyone else and still thinks their approach is the best, yet evidence to the contrary.
It’s simply how times change and people with it. Knowledge is lost when people move on and the reasons why certain decisions were made are not transferred.
At any rate, I imagine people at Google are trying to figure out why there is such a negative opinion on their search results lately.
Why should they care? They're too big to fail…
Google controls almost the whole end-user realm through Chrome & clones, and Android, the dominating end-user OS by a wide margin.
At the same time end-users are completely helpless and can't do anything against Googles liking because they don't understand anything about IT tech.
Computers are black magic to most people so they're trapped. This never changed! Especially millennials and gen-z are completely clueless as they didn't had the chance to use personal computers ever, where you had at lest some control over the device and needed to know at least some basics about its inner working. All the younger people know are the tightly sealed black-box devices you don't have any control over, called mobiles, which are fully operated by big-tech. Google search + Android apps are "the internet" for most people. They mostly don't even know there is something else beyond that, so Google can do whatever they want, and this will have exactly zero consequences for them by now.
Google's move regarding rolling out "browser DRM", the next "trusted computing" initiative, regardless of what anybody thinks about is is very telling.
Now they will violently reap the fruits of their monopoly, and likely nobody will be able to stop them in the next decade. People where warned about the consequences of this monopoly for many many years. Nobody cared. Now it's payout day for Google.
When I pointed out that it's not the responsibility of the one answering to search for dupes, but for the one asking, I was told that I should still invest the time or otherwise don't answer at all.
SO will remain as a library for unseen problems, which is what's supposed to be good for.
Stack Overflow has a lot of stridently opinionated jerks contributing to it, and if I can just ask ChatGPT a question and get an answer that works rather than having to deal with being belittled by those people, then I'm probably having a much better day as a result.
Remember all users are moderators. There are some explicit moderators but they don"t close or downvote often they deal with other problems - or on smaller sites just use normal user powers to vote and close.
The fault finding culture of SO is toxic.
The aim of SO is to provide answers to a question.
You do not want many questions with the same answer as if you have a new answer or a comment on this duplicate answer you need to then add it to all the questions. Thus we want to collapse all these multiple questions into one.
Also the person who I was replying to did noit seem to understand that they were a moderator, moderators are not a separate set of people to users.
However SO Inc wants money and so more users.
You can"t expect the answerer to answer each individual question which is the same, especially as the answer might have been given over 10 yuears ago?
What do you expect a good answerer to do with a question asked for the hundredth time?
I had a hard moment on the gamedev stackexchange where I was stuck trying to learn how to do something in OpenGL. A moderator immediately closed my question as a duplicate because there was a similar question about OpenGL ES, which is a (related but) different API. I tried to plead my case, but was shut down.
Shortly after that, I gave up on the game I'd been working on for a couple years. The mod's decision contributed to that.
I felt stuck by a wall between me and answers to some of my game programming questions. Over-moderation is more than an inconvenience. It can destroy the ability of users to get things done.
Looks like Google started to prioritise ads even more than actual useful results is what changed mostly.
I do typically self-answer if I figure it out, but you know, if I'm going to be ignored maybe it should be a github issue so I can get the sweet zero replies and that juicy 90 day auto-close from inactivity.
My friend, a much less skilled developer, was much better at the craft of SO. I just lurk/solve my own issues
I have only had great experience with SO, but regardless, don't get dissuaded by those with less than mediocre tact.
The quick way to accumulate reputation is through bounties. However, it's very much a lottery: bountied questions are often about ultra-specialised niche topics. You may need to hunt for a long time to find something you actually know something about.
Thus I don"t understand your issue. This is a XY-problem :) I know enough about the subject to know that your issue is not the actual issue as anyone can answer, if you had issues then there is something else.
1. it takes time to write a good answer;
2. fast copy-paste answers, often wrong, get upvoted anyway;
3. picking the accepted answer, out of many, is done by the person who asked it... who usually doesn't have the knowledge to do this!
4. also many times no answer is accepted, like the person who asked it stopped bothering.
And the hard blowback for that old post on how to game the SO point race, instead of SO trying to figure out how to fix their broken system.
Just like SO was a move out of Experts Exchange and others, this is a good time for the community to start a post-SO Q&A site.
Answer a difficult question on a barely documented part of software (e.g. low-level) and you'll get a couple of votes, at most. And you're lucky if the answer gets accepted.
There are a few unsung heroes on certain hard/obscure SO tags. They dedicate a lot of time and get little reward. Whatever follows SO should find a way to fix this.
- google used to return really relevant results for SO, and it stopped doing so at some point a while ago
- moderation on SO has gotten progressively more horrible. can’t tell you how many times I found the exact, bizarre question I was asking only to see one comment trying to answer it and then a mod aggressively shutting it down for not being “on topic” enough or whatever.
- because of the previous bullet, oftentimes the best answer is buried in comments and has very negative feedback despite answering the exact question
Due to a combination of these things, filtering against the noise for what I wanted became increasingly more difficult and often the solution to my problem was easier found searching github comments or random blogs.
OP: “how can I accomplish X doing Y?”
top voted response: “I wouldnt do X by Y. Instead, I’d consider Z. It’ll do everything X by Y would do, plus these other things.”
buried sub comment to top voted response: “that’s not what he asked, to do X by Y (exact solution provided)”
then 10 comments blasting the guy who actually answered the question for not doing it the way the top voted answer did it.
Case in point: https://stackoverflow.com/q/19622198
What's up with the "man Y" aversion towards the usage of examples in the man pages? I don't expect some behemoth like ffmpeg to have a billion examples in the man files, but damn, every other reasonably sized CLI app would be so much easier to use, if you had just a dozen or two of examples at the end of the man file.
tldr ffmpeg
Sometimes, the original poster is not mistaken. Especially when an expert asks a question, there's a reason for it. Assuming the expert to be a beginner who hasn't tried easier solutions is degrading, and forces experts off the site.
Actively insulting your userbase, especially your expert-level userbase, is a bad idea. It leads to StackOverflow falling and collapsing over the years.
IMHO Copilot Chat has picked up on these bad habits.
I literally cannot remember when last SO was useful time, due to the false positive identification of the XY problem.
At least with ChatGPT, it's much faster to get it to answer the question asked and not the question that was not asked.
the XY problem-problem
I never saw this term before, but it is perfect. So many times I have experienced it when asking questions in mature tech domains. It is so frustrating. Frequently, my question is asked poorly from the view of an expert, so they sweep it aside as not a real problem. Only after getting help (from comments) to provide more info or improve the writing, does the expert suddenly agree it is an issue.This is also why I no longer waste my time raising bug tickets for open source projects. You just get shouted down and feel terrible about yourself. I raised many "WONTFIX" bugs in my career against open source projects. What a waste of my time and a harm to my self esteem.
So I do know this problem (problem-problem) very well and I've since stopped asking or answering on SO.
This is called the XY problem and is extremely common on tech related forums and mailing lists.
"I'm trying to learn foo.js and attemt to do so by porting Tetris to it. I know foo.js is a terrible option for a game. So. How can I write to the canvas from foo.js?"
Is often answered properly, if only because to many it's a nice puzzle.
"How can I write to the canvas fro foo.js?" Is different in that it will attract a lot of people explaining that foo.js deliberately did not allow writing to the canvas, because Z"
This may help less experienced engineers not understanding their problem set, but for a more experienced engineer it's absolutely obnoxious to think of all the ways to defend my question so I don't have to deal with a rush of "oh but you should do this instead" answers getting up voted that don't actually answer my question, then asking to accept the answer.
To one set of users it's possibly helpful, to another it's useless if not also condescending, often condescending to both sets.
The bottom line is now, SO is so filled with these types of responses, I can't expect to get a very specific question answered, which is really the only reason I'd ask a question in the first place, so why use it?
There are plenty chat groups now via Slack and discord in my field where I can get much more direct answers. People aren't worried about getting down voted for a bad question, people aren't giving low quality answers to boost their points. So for me, SO is practically dead except for the occasional obscure error message that I can query for there.
My experience is that it's not so much a defensive argument, but context. My example was poor in that it could be misread as a defensive argument, sorry about that.
I meant it to show how adding some context changes the question. Because, in programming, it is all about context. AKA that "It Depends" meme.
Comes up all the time when people ask how to do things in bash/sh. I know there are better tools for the job, but this is the one I have.
Explaining what the XY problem is to people who are telling you about it's high false positive identification is, itself, an XY misdiagnosis.
Your reply is an example of what people are complaining about - you are addressing the issue you wished was asked, not addressing the issue you were presented with.
But those are correctly diagnosed XY problems. No one is complaining about those.
My parent was told the issue is too many incorrectly identified XY problems, and responded with an explanation of what the XY problem is.
That is the example of a misdiagnosed XY problem, which was kinda my point. This sort of behaviour makes the actual experts leave the site in droves.
If, when answering a question, one where to discard the answer the minute they write "Why would you want to do this?", you'd get much fewer incorrectly diagnosed XY problems.
As I said in a different thread, ChatGPT sometimes does this as well, but at least with ChatGPT, when it is answering a question that was never asked, it doesn't also act like a condescending jackass. There are no "Why would you want to do this?" type of questions.
But on second consideration, I suppose you would not, either, and I suppose you are specific'ly talking about responses which, each taken as a whole, are easily interpreted as some form of hostile.
I responded to the person above me when there were literally two other comments in this thread.
> Your reply is an example of what people are complaining about
Defining something is not an example of an XY problem.
> - you are addressing the issue you wished was asked, not addressing the issue you were presented with.
As I am not much of a programmer but work on electronics and computer hardware I deal with different types of people than would be on SO, so I am not addressing anything but my own experiences.
And then to ice the cake you find the question because your question has been marked as a duplicate of the older question where they answered using unix binary tools… and you specifically asked about doing something in “pure shell script” or something similar to that phrase.
Stack overflow is fundamentally a system design that breaks down at scale due to misalignment of incentives that are necessary for it to work well at smaller scales (as can be seen in the successful operation of various smaller Stack Exchange sites for various topics such as Law, Aviation, Physics, etc)
I must admit that I don"t buy in to that philosophy and like using one tool so for scripting I would do all in python, so I would not be answering that question.
Many Linux users think that their way is the only and that means bash as the shell and many others like GNU coreutils, gcc etc. I am an MacOS user and my professional career includes several non Linux Unixes so I know bash is not the only shell - try csh for fun - which is partly why I use python or previously perl as they are the same on all machines
You may know, but someone else asking the exact same question may not.
So sorry then, but listing every constraint and design descision it is.
You don't need to do that, a simple "I know what the XY problem is, and this isn't it" prefixed to every question you ask should be enough to stop the race to tell you all about the XY problem.
I mean, at this point it's clear that more people know about the XY problem than people ho don't.
True, but that logic goes both ways: Unless told otherwise, whoever reads the question, isn't required to assume that there is a constraint.
If I get asked how to water plants with a sieve, without getting told why getting a watering can is impossible, "You don't, use a watering can", is a perfectly acceptable answer.
Especially when the question is aked in a question archive, rather than a help-forum.
If specific constraints apply to a question, then they should be a part of the question.
But I think it will always be up to the user of SO (not the poster or answerers) to make the real judgement on what is useful.
Often I think SO is useful to use as a bunch of puzzles folks solved. You gotta decide if they are relevant.
SO is at its best when it’s actual error debugging IMO. When you google some specific error whoever else has the similar error it’s right there. I feel like GitHub is replacing this more and more though - I often get the GitHub ranked specific error higher than Stackoverflow these days. Usually you get better discussions on the GitHub issues too, for a multitude of reasons. Two off the top of my head:
1. all of the people working on the stuff related to the issue are very close by
2. the moderation is not nearly as heavy handed as SO.
ChatGPT is also much better than SO as well if you can give it enough context and the thing you are working on wasn’t built on stuff released after 2021
I also really like Stackoverflow for current event type stuff, like black swan type events. One recent example is when google’s Paris data center was on fire and infra guys were helping each other out trying to get systems online.
All of this combined means that StackOverflow the forum is probably on its way out though. They made the mistake of taking VC money and the model hasn’t really proven profitable so they have really made some poor decisions to please the vc overlords.
I won’t miss Stackoverflow much other than nostalgia unfortunately - better alternatives have arrived. Seeing the decline of all of the other Stackexchange sites kind of sucks though. There aren’t better alternatives for many of those
Both sides can identify ChatGPT answers as being wrong. The question is how can they be deleted. The moderators say they can delete a lot by manual inspection. SO say that AI tools were deleting wrong ones.
Which is more than Google, or a previously answered StackOverflow question, can say.
But then, it’ll probably greatly vary depending on the problem you are facing.
Bots closing issues because someone doesn't spam the page. Closing as duplicate of (non related bug). A slew of random solutions that are only tangentially related and don't really solve the problem.
I’m sure if you’re twitter Kafka is actually a better solution, for everyone else, it isn’t.
Want to host videos on a laptop (which has a big SSD) and stream them to a Pi (which is attached to a big screen) over a LAN? Hey, here's a post about how to host videos on a Pi and stream them to a laptop! Upvote and share! My point is, you don't even have to be trying to do something all that strange for people to apply the XY Problem logic and refuse to help you.
(Solution: NFS mount and a patient understanding that the Pi cannot play certain kinds of video, so you'll need to transcode some of them first. See? Nothing bizarre, but surprisingly outside-the-box given what I could find online at the time.)
One common example I've run into lately, as I've been reading about state machines, is people asking how to implement a simple react component as a state machine, and others objecting to the premise of the question since using a state machine for a simple react component is obviously a bad idea.
It’s fine, I think, to answer a question and then suggest a better method. It’s presumptuous in the extreme to dismiss a question with some pseudoacademic neologism.
IOW, any time you think you have spotted an XY problem, you're probably wrong.
And that's the problem with SO moderators and regulars. They classify everything as an XY problem because it allows them to answer the question they know the answer to rather than answer the question that was asked.
The thing people don’t get is: when you answer on SO you’re not answering that poster. You’re answering anyone who will ever have this question. It’s quite arrogant to assume it will be an XY for every single person forever more.
The proper way to answer is to answer the question exactly as ask and then insert your “but you probably should be doing Y instead” at the end.
Doing things the right way is BETTER.
If you can't, you should add a bit to your question saying "I know the standard way is to do Y, not X, but because of reason Z I can't do it."
Most readers will be able to use the "right way".
My experience with less experienced developers is that they ask the exact question as that is where they have stuck but they are ignorant of the better ways.
I do tend to answer differently depending on the questioners reputation. If they have a higher rep then I can assume they know what they are doing.
Because you're answering every person who ever will ask, a lot of the people who pass through your question & answer will be people who don't know the difference between the right way and the wrong way. If they want to know how to do something the wrong way, because they don't know what the right way is, an answer that simply tells them how is a bad resource.
It's not enough to tag caveats onto such dangerous answers, because people can't read. Instead, newbies should have to overcome a sufficient amount of opposition to filter out those who don't know why they're doing what they want to do, and the rest can make the little effort of being very explicit about why they want to do something the wrong way.
Can't be done, will be marked as duplicate.
Then Googling for doing X with Y gets you a bunch of closed questions and a labyrinth of links all leading to a question that was answered 10 years ago on a different software version where Z possibly was the right way to do it but now isn't.
And of course there's no way to reopen the question because it has been closed by a level 15 Magister Templi moderator and a lowly level 3 apprentice moderator like yourself needs to either answer 146 more questions or moderate 192 other questions to clear enough arbitrary hurdles to achieve holy question reopening powers.
And there's possibly an appeals process but that involves recruiting 13 moderators who you have to convince to give this question special treatment and declare that one of their number of sacred moderators made a mistake.
Some of it is that SO has gamified shitting on and suppressing the question/asker instead of gamified providing the answer, and built a culture of toxicity that tolerates the abuse of the tools in this fashion.
And when the CEO asked them to tone it down maybe 5 years ago they basically did a collective “am I so out of touch? no, it’s the askers who are wrong”. Extremely funny to read the meta responses to that at the time.
https://stackoverflow.blog/2018/04/26/stack-overflow-isnt-ve...
https://news.ycombinator.com/item?id=16934942
(admittedly "women and people who don't speak english well are particularly unlikely to adopt to the pedantic neckbeard culture we've built" is a spicy take for your average SO'er, or wikipedian, but it's also not actually a wrong one either. SO's culture problems probably do disproportionately chase away users with marginal engagement, nobody likes putting up with formalized neckbeard culture and those users have absolutely encountered it before and absolutely have an aversion/revulsion to entering yet another online neckbeard nest. I think this is a case of “he’s probably right but the medicine would have gone down better with the manchildren if he hadn’t mentioned women and minorities”, and he’s also right that those issues have continued to bury SO over the last 5 years.)
Then you have to do two things in your answer:
1. Correctly answer the question as asked.
2. Add your opinion about the "right way" to do it.
If you only do #2, you are failing "every person who ever will ask."
[0]: https://www.joelonsoftware.com/2000/04/26/designing-for-peop...
You know that's not your responsibility. If some newbie makes a mistake, that's their responsibility (and a learning experience for them).
And frankly, I think you greatly overestimate how valuable and essential your non-responsive "you're asking the wrong question" answer is.
> https://www.joelonsoftware.com/2000/04/26/designing-for-peop...
That link is about users. You're misapplying its lesson if you're using it to justify not answering a developer's development question.
Quit coming up with excuses for not answering the question.
Why do you think "users" is an inaccurate description of the role question askers have on a developer Q&A board?
Put another way - when was the last time you used a development tool, or a library, or some other resource, and sat down to read the full documentation of it? I would posit that that's very rare as an activity, even for developers who need to develop a deep understanding of what they're using.
It's much more common to learn by doing, and the limit of that learning is very often what the developer can't do. Answers which easily enable developers to do something are overwhelmingly likely to lead to developers doing that thing - much in the same way that a long page of library documentation which gives an example is likely to lead to developers repeating that example, even if at the end of the docs, there's a little caveat saying that you shouldn't follow the example for so-and-so reason.
>If some newbie makes a mistake, that's their responsibility (and a learning experience for them).
But is it a good experience? Sure, maybe they'll learn that they always have to read the whole answer before they use any part of it. But we sensibly have abandoned this no-guardrails approach to teaching in almost every arena where it's been used, because it's not really suited to the way people do things in real life - and in real life, people often end up affecting others with their mistakes.
Does junior developer who learns how to glue SQL strings together in their favourite programming language, and makes the "small mistake" of not learning anything about SQL injection in the process, benefit from the learning experience when they cause a data leak? Do their customers? Or should the learning resources they access maybe use the pedagogical tools available to make sure those kinds of mistakes are really hard to make, even if it occasionally inconveniences a seasoned pro?
The are a lot of different kinds of "users," and I think the kind of thinking in that article is totally inappropriate when applied to developer Q&A board.
To be perfectly blunt: the result if what you're advocating is to condescendingly treat experienced people as newbies so dumb that their question should not be answered, because you think they're so dumb the real answer might distract them from the lecture you want to condescendingly give them.
People like that are super annoying and almost always unhelpful.
Every single fucking question I ask on SO has some lazy condescending dude chiming in to answer the easy question he thinks I should have asked, after he totally failed to understand the constraints that made my question hard. Of course, lazy condescending dude always thinks he knows better.
Here, I'll coin a name for a problem I see much more often, which is called the XX problem:
1. User has problem X, and asks for a solution for it.
2. People viewing the question decide that the original user actually has problem Y.
3. Those people tell the original user that they actually want to solve problem Y, condescendingly flame the original user for not asking about problem Y, and if they have the power to do so, edit the original user's question to be asking about problem Y.
4. Those people use poorly-thought-out pop-social-psychology to justify their shitty behavior.
5. The original user still doesn't have a solution to problem X, and they really needed a solution to problem X all along.
It is not without irony that examples of the XX problem are sometimes also examples of mansplaining.
You realize that it's dismissive and condescending to ignore the problem I'm describing and respond with an assumption that I'm unaware of the tone of my post, right? Pot, meet kettle.
Because you're propagating the idea that you (or anyone asking questions) knows better what a person asking needs than the person asking does. Telling people you know what they need better than they do is pretty close to the definition of condescending.
People are emotional because nearly everyone who asks a question on the internet has to deal with people telling them "You don't actually want an answer to your question, you want an answer to this other question." By boosting the signal of the horrible XY problem idea, you're contributing to that problem.
I'm not saying this kind of miscommunication never happens, but the opposite is actually far more common.
Unfortunately I had no idea about the perverse incentives for question answering and the terrible moderator practices on SO, so I walked into a minefield giving an answer here that I thought was helpful based on my experience as a hands on technician working directly with people -- but it turned out to be a lightning rod for people's frustrations on these issues.
But apparently I can't comment as a new user so I guess the discussion will just be "because calculus classes use matrices in examples".
[0] https://matheducators.stackexchange.com/questions/26417/why-...
What's the name for the "I don't want to use a sledgehammer to solve a problem that should be solvable with a screwdriver" problem?
Or in this case, the "spinning up 300 lines of code to integrate an XML parser vs. a dozen lines of code based on regexes" problem. For reasons that are unclear to me, XML parser libraries tend to be painfully difficult to use (speaking from personal experience with 4 different XML parser libraries).
I don't think it's a surprise to anyone involved that an XML parser is going to solve the problem.
1. To do Q that way you would to <solution, or at least pointers to how to find the solution elsewhere> but you will likely it far more efficient/easy/whatever to do <something else> instead.
2. Q can be achieved far more efficient/easy/whatever by doing <alternative>, but if you are stuck with using Y then try <solution or pointers as above>.
Of course this relies on you correctly deriving that they are trying to achieve Q, or them explicitly stating the fact. Maybe instead they are trying to get to K.
It probably won't last forever without people generating new answers somewhere else, but it answers a lot of things correctly, and the things it gets wrong are easy enough to verify.
I really hope ChatGPT causes StackOverflow to change.
It already has, and not for the better. SO is currently awash with nonsensical answers that are clearly the result of feeding the question into ChatGPT. Not to mention questions like this: https://stackoverflow.com/questions/76748781/how-pythons-bui...
Yes, there are plenty of rude replies, but sometimes it helps to assume what somebody really means is they failed to do a context switch. Exchanging more details can help turn a 'madness' into a method. Even if they were outright rude, doing this can lead to an answer that might not otherwise be given (even though that answer may be provided by someone else).
The user asks a question that can be answered quite easily, and dozens of people post answers claiming that this is the wrong way to do it and that they should use some other tech to solve the problem.
Some people on Stack Overflow care more about showing off how smart they are rather than answering questions, and I think the point system attracts these people.
[1]: https://stackoverflow.com/questions/1732348/regex-match-open...
I actually think SO is a great site and resource, but I also think a lot of that is despite the bitter old timers in the community, not because of them.
As someone who was a newbie at the time when it was posted, who was looking for a way to parse HTML, I took away that it's just really the wrong way to go about it. I didn't feel crapped on at all.
But the asker is very clearly a newbie. The question does not contain further context. The asker's suggestion is wrong (I think). And we've all worked with junior engineers who try to use the wrong tool.
The answer is a whimsical way of making an appropriate suggestion in this inferred context.
Also, to be fair, I think it's not mathematically impossible to use dark regex magic (with look-behinds and such) to parse HTML, but that's a discussion for another day...
If you scroll down long enough, you will see answers explaining that. But they arent upvoted as the answers suggesting the questioner is an idiot.
I and I think many other are sad that S.O. removes many serious work related questions (have lost count of how many times I saw the perfect question with the perfect answer, with a note that this isn't what Stack Overflow is made for and these questions only exist for historical reasons).
As mentioned elsewhere, this is an old question and both the kind of question and answer wouldn't be allowed these days.
However, I also fundamentally disagree that questioning the assumptions in a question is unhelpful. You want to solve a problem, find an approach and want help because you have problems with the approach? What if the approach you took _is_ wrong? Its very helpful especially for advanced beginners or at the intermediate level, to be given a different way of solving the problem even if that is not what you asked.
It depends on context if this is just pedantry or genuinely helpful. The best answers I found start with answering the question that was stated, but then proceed in showing how the problem behind the question can also or better be solved.
I'm fine with indirect, "question-the-premise" comments, but they should be posted as comments, not answers, because they are in fact not answers.
Stack Overflow isn't a site for beginners, it's for "professionals". At least, that's what all the Stack Overflow defenders tell me every time I criticize the snarkiness, rudeness and patronizing manner of many answers/comments you receive on Stack Overflow.
Jesus this is cringe
So the answer is actually on topic.
It may be a bit… stylistically unusual, but I think I came away with a pretty good idea of what the answerer thinks.
Yes, and they are factually correct in doing so. The correct answer to the question: "How do I use a hammer to tighten a screw" isn't a lengthy description of how to perform some magic with a hammer. The correct answer is: "You don't. Use a screwdriver."
HTML isn't parseable with regex. The various answers under the question explain in great detail why [1] that is the case.
SO isn't a help forum, it's a question archive. The purpose of an answer isn't to solve one guys specific question, but to provide an answer that is useful to all people who ever stumble upon this question.
Yes, and answers on that very page, with lots of upvotes, do exactly that. People looking for answers online, can reasonably be expected to scroll down a page with results.
My car broke down in the middle of the desert because of a screw that came loose and all I have is a hammer. You have just condemned me to death because you assume you know better.
Poster is not asking for this. He is asking how to parse a specific subset of HTML. And it is demonstrably parseable.
> The correct answer to the question: "How do I use a hammer to tighten a screw" isn't a lengthy description of how to perform some magic with a hammer. The correct answer is: "You don't. Use a screwdriver."
It is not the appropriate way to tighten a screw, but there likely is a correct way to do it with a hammer.
It is fine to point out that there are better ways to parse that HTML, but it is not wrong to do it with regexes.
Sorry to be blunt, but having coworkers like you make the job really annoying. I'm not a newbie, but a seasoned programmer. If I'm asking a question and am in a domain with a fair amount of experience, don't give me patronizing answers.
Poster is not the one answers are for. Answers are for everyone who stumbles upon this question in the future, and the general topic of the question is very much about parsing some HTML with regex.
Again: SO != Help FOrum
> but there likely is a correct way to do it with a hammer.
No, there isn't. Because the correct way is to use a screwdriver. There is certainly a way to do it with a hammer, same as there is a way to write a webserver in brainf__k. Doesn't mean that way is good or should be done.
> Sorry to be blunt, but having coworkers like you make the job really annoying.
Bluntness is fine. I will be blunt as well: Having to fix code full of hammers used to tighten screws is a lot more annoying than having colleagues who try to prevent a codebase full of hammers in the first place.
So, no, the top 3 answers are all bullshit.
The questions isn’t about parsing, the question is about recognizing a token.
So here is the reason: the top voted answers are wrong.
"Some" is an understatement.
My favorite example, from "TeX stackexchange" :
https://tex.stackexchange.com/questions/100574/really-wide-h...
Quote from the actual answer (which is not the bogus accepted one):
The question wasn't "should it be done?" But, for the same reason men climb mountains, "could it be done?" The answer, [...] is yes. Thus, we introduce [...]
> How do I shoot myself in the foot? I’m pointing the gun at my foot and pulling the trigger, but nothing is happening.
Is it:
> A common reason for guns not firing is that they aren’t loaded. Try loading the gun and trying again.
…or is it:
> Woah, hang on a sec! What are you trying to achieve exactly? It seems like you are doing something very wrong here. I’m sure there’s a better way to do whatever it is you are trying to do.
There are many technical questions that give the very strong impression that somebody is asking how to shoot themselves in the foot. It’s not responsible to blindly answer the question regardless of the consequences. Yes, people sometimes overcorrect for this which can be annoying, but they are only trying to steer newbies away from shooting themselves in the foot.
But it's hard! On the one hand, you don't want to teach people to do the wrong thing (especially not on SO, where answers are often copied wholesale and pasted into production code), but also not answering the question as asked usually doesn't help the OP at all.
If it's not clear what they're trying to achieve I usually leave a comment asking for clarification rather than answering, though.
Sure, but some of the stuff I've looked up in the past is answered with helpful comments like "You won't need to do this, your CA will do this for you" or "this is handled by your certificate verification stack, you don't need to be involved in this stuff".
Well, thanks, but I'm actually implementing both of those things right now and I'm having an annoying issue with figuring out how the API for the (very popular but IMHO poorly documented) library I'm using for part of it hangs together.
> Yes, people sometimes overcorrect for this which can be annoying, but they are only trying to steer newbies away from shooting themselves in the foot.
I think sometimes its because comparatively the respondents are newbies, and are repeating received wisdom.
Thus, the first response is the correct one.
On preventing people from doing what they intend to do...I think the main issue is the world is a large place, people face many situations, and most advices given in good faith only have an extremely limited view of what people might need to do. And it's kinda awkward to ask for someone's whole life to decide if they are right in wanting to straight rename a table column on their live production DB.
Something could be a bad idea 99% of the time. But that leaves millions of people in the 1% for whom it's the best course of action.
When I find a straight and relevant answer that is great, I like browsing SO, but I wouldn't dream having a question just to find myself in a word duel with some very eager contributor steering me into his favourite domain. I want to get ahead with a task not being in the quest of finding The Ultimate Solution or having a social life over interesting and related topics.
This is one I ran into recently:
https://stackoverflow.com/questions/30196175/const-methods-i...
The question is very clearly formed. The accepted (and currently top) answer does a good job. But at one point the other answer which is badly worded and confusing was at the top. And I don't think it should ever get the top position of the answers.
It's ok, you deserve the hate. After all, you're asking the wrong questions.
They want questions that are "textbook-worthy" or possibly "encyclopedia worthy". They're on record as having said this officially. If programming and technology is just messier than that, if it's more complicated than that, and you still have questions that don't fit because of it... well, fuck you. They don't exist to answer your questions, they exist to answer great questions that they can use to build up this pretty little sight that now has some purpose other than whatever it was you thought you were using it for.
The Z guy, he's playing their game. He's awesome. You, you need to be punished until you comply.
I wanted some information on figuring out what texture serialisation was supported on a given client for a WebGL app. I needed to know this because I was optimising the client and had to deal with very large textures (it was an AR / augmented reality context).
Cue a barrage of comments along the lines of "you shouldn't need to do this" "the abstraction means you can simply assume the GPU has infinite texture memory" (!) "just provide all the formats and let the GPU bridge figure it out". Then the question had a downvote and that was that, cast to the bottom of the pile.
It seemed to me everyone responding to my question had this assumption "questioners are morons". A moron asking about texture serialisation is a paradox, ergo the question must have a faulty premise.
I wanted good ways to deal with pointers to 2D arrays in C, which aren’t hard but you need to remember that [] has precedence over *, so (*pointer)[x][y] is needed to deference. It’s not the hard to mess it up. Ultimately GPT had an OK answer but… it had that immediately without searching and the didn’t have to craft some well written example, it picked up my dirty explanation no problem.
I have the 2D array typedef’ed now, but it’s still confusing to read and hard to work with. I’ll search tomorrow on it.
Also, in C++, you can set your clock to someone commenting that you should just use smart pointers. It doesn't matter if the question is entirely unrelated.
As someone who spent a decade helping people answering code questions on Flashkit and later SO, I find the SO community and moderation now to be so off-putting that I avoid asking anything there if I can. I still give answers sometimes, but I'm much less likely to be on the site at all.
Rather than ask it to write a full ansible job, try asking it the question you would ask on Stack Overflow (i.e. you probably wouldn't ask Stack Overflow to write the full job for you).
Maybe they should expire comments or remove them alltogether.
Or something vaguely like that?
Doesn't hacker news also do this?
Edit: It seems like there’s a wiggle room, but still terrible. https://news.ycombinator.com/item?id=28293237
Reddit is actually a lot better in this regard
I wrote paragraphs only to be closed by mods as “opinion based” or whatever, so I had to turn them into blog posts, and they got picked up by popular publishing channels, like this one: https://medium.com/hackernoon/are-interfaces-code-smell-bd19...
Of course I care about how my work and my relationship with it is treated. Why don’t you?
And I wasn't talking about the issue of mods deleting entire answered questions. I was talking about single contributions, and mods deleting things is the exact opposite of what my post was addressing!
On the other side, that compiler post is fun but it's not some big effortful thing. Caring about it some makes sense, but I wouldn't care about it that much.
It wasn't a blog post on the site; I extended it to a blog post after mods closed it and flagged it for deletion. As a top 0.5% or whatever contributor of SO, believe me, I know what the site is for.
> but it's not some big effortful thing
It doesn't have to be. It's my thing. It sounds like me, it carries my spirit. Nobody else would have written it then. It reflects a period of my life, how I perceived things; and more importantly, its continued existence has an impact on me regardless how it's licensed. Say, if my writing style, or my choice of words, or my tone were associated with a traumatic event in my past, SO's insistence on keeping it up would be explicitly abusive, don't you agree?
Even if no such event had occured, SO's or HN's insistence on keeping my content online against my wishes is also abusive; it deprives me of control over my thoughts and my words. Is it legal? Yes, 100%. But, is it ethical? No, I don't think so. And for what? For keeping the answer for "how can I simulate a click on a DOM element?" online, as if that problem immediately becomes "unsolveable" when that comment, heck, the whole SO web site goes down. What a pretentious excuse to keep your income stream steady.
> I wouldn't care about it that much
You would if you were me.
No, I'd say your trauma is causing you to make an unreasonable demand. The writing style in a short technical post existing somewhere should not be harmful.
> SO's or HN's insistence on keeping my content online against my wishes is also abusive; it deprives me of control over my thoughts and my words. Is it legal? Yes, 100%. But, is it ethical? No, I don't think so. And for what?
It's supposed to be a collaboration, especially SO, and keeping things intact is important for that.
Just yesterday I found a helpful guide on reddit where half the posts were "." It's pretty clear how a site where the primary purpose is guiding people would do a worse job if it worked that way.
> And for what? For keeping the answer for "how can I simulate a click on a DOM element?" online, as if that problem immediately becomes "unsolveable" when that comment, heck, the whole SO web site goes down. What a pretentious excuse to keep your income stream steady.
If it would affect the income stream, then it's something that makes the site bad for users. You can't argue both sides of that at the same time.
> As a top 0.5% or whatever contributor of SO, believe me, I know what the site is for.
Maybe? I don't think most contributors would be anywhere near as upset about an inability to delete posts.
SO wants text like you gave them, I don't think they want the level of emotional investment in those pieces of text. (They might want emotional investment into the site itself, but that's a different thing.)
Most contributors of any platform wouldn't be upset about anything. Conformance of the majority doesn't imply the rightfulness of the action.
But many people are upset about it, rightfully so; see for yourself: 1. https://www.google.com/search?q=why+can%27t+i+delete+my+cont... 2. https://www.google.com/search?q=why+can%27t+i+delete+my+cont...
> I don't think they want the level of emotional investment in those pieces of text
Of course they don't. They want you to be a free ChatGPT as much as possible. The less human you are, the better for them. That doesn't mean that what they are doing is okay or harmless.
This is a lot better than almost every other comparable site with user-created content. In most cases the company has the rights there and can do whatever they want with it, and users have no right to reuse the content of the site.
E.g. Python 2 answers aren't relevant anymore, but still there.
I see "answer as comment" in my most mature tech subjects because people are afraid of downvotes. In my experience, the most mature tech subjects are carefully guarded by a small community of very unfriendly "fastest gun in the West" types. The C++ community is incredibly unwelcoming and negative towards most questions. Note: Comments can only receive upvotes. (In theory, comments can be flagged as off topic, or offensive, but assume -- for my commentary here -- they are not.) Examples of more mature tech subjects: Python, Ruby, ASP.NET, Java (language/foundation libraries), C# (language/founding libraries), C (but not yet C++!), Win32, etc.
For less mature tech subjects (or faster moving), you see many more answers. Example of less mature tech subjects: Python AI/ML libraries, Java/C# open source libraries (Spring, etc.), C++, Qt, Gtk+, Zig, Swift, Golang.
People are afraid of downvotes because people will throw out downvotes without even reading your answer, and once you get the first downvote, everyone else will pile on, also without reading your answer.
Also, the site itself - the ethos of the owners; I used to have an account, became disenchanted with the mods, decided to delete my content, close the account.
I then discovered;
1. You cannot delete answers which have been accepted.
2. The mechanism for deleting your own content allows you to delete a max of something like five replies per day.
I then took the time to look at the T&C, and SO as I remember it simply make all of your work their property.I deleted everything I could, over the course of a significant number of days, and left.
Your answers will most likely stay, just anonymized.
Do note that your posts on SO are published with a CC-BY-SA 4.0 license:
https://stackoverflow.com/help/licensing
so regardless of who owns them, they can be reproduced, and probably are. So in many senses there is no deleting them.
Also, don't understand why you wanted to delete. Despite its problematic policies, SO is an important public resource and I assume so were your answers on it...
So remember it was never your property.
Behaviour of users on SO has always been, well, as all the complaints here suggest. It's never really bothered me that the top answer's someone going off on their favourite approach, most questions have multiple answers, and I've learned a lot about coding from reading the top few and trying to understand where the answers are coming from.
The discussion's as important as the answers sometimes, and that's why a potential collapse is a problem.
> - google used to return really relevant results for SO, and it stopped doing so at some point a while ago
SO might be horrible now, but it still holds years worth of answers that were just fine a few years ago - so why aren't they showing up now? Google's current recommendation of going to w3schools or - even worse - geeks4geeks or any other content farm is and always will be worse than stackoverflow. I don't have a clue what their algorithm is doing but it's surely trying to kill Google search as fast as possible.
Another joke is the fact that searching for "[language] [symbol]" also brings me to these content farms instead of the documentation. You seriously can't find useful anything these days using Google.
I would say at this point you could just use wikipedia as your default search engine.
Only way below this block (which takes about 120% of my whole screen's height) come the "organic" results, that aren't great, but probably match what Google assumed I wanted to see.
The business reasons why Google doesn't take steps to remove the bad content and make their product pleasant to use again is so far from my understanding it might well be aliens running the company for all I know.
Thinking of it, it would be an interesting test to compare the ranking of two similar sites, one with google ads, another with ads from another provider. Might be good evidence for antitrust litigation. But what do you do if they just prefer sites with more ads? Because due to their market position, that benefits them, but it isn't anti-competitive against other ad-pushers.
Just managers doing what they get more money for or devs hunting promotions by increasing ad revenue by 0.01% in the short term one sting at a time.
It's really good and IMO more than worth the price.
[0] https://kagi.com
Some people argue that Google possibly can't win the fight vs spam sites, but obviously it works perfectly fine manually blacklisting them.
It takes time to build ranking.
The underlying reason is probably that the spam sites use Google Ads (revenue which is tied to 1000s of PMs and managers bonuses) and that Google as an org is deeply dysfunctional at this point.
And not in a generic “this kind of content is bad for our users” but very specifically “site x.com is bad for our users.”
> Yeah, but then they're editorializing.
> And not in a generic “this kind of content is bad for our users” but very specifically “site x.com is bad for our users.”
Except that shouldn't be a problem because I'm pretty sure Google already blacklists domains.
These scam sites load megabytes of junk, load slowly, have text interpersed with ads and modals that render right on top of them, if you open devtools you'll often see pages of warnings about deprecations and/or invalid html, and despite having the same scraped text always score higher on Google.
Only if you're not googlebot. The crawler sees a much nicer site.
"If you operate a paywall or a content-gating mechanism, we don't consider this to be cloaking if Google can see the full content of what's behind the paywall just like any person who has access to the gated material and if you follow our Flexible Sampling general guidance."
I wonder if they just gave up
If I look at Google's guidelines, my articles follow all of them: in-depth, well-researched, demonstrating personal experience, better than other articles appearing in the search results. And yet, they were "penalized" by this update for who knows what reason.
I looked into it and some other websites benefited from the update, so who knows what changes they made and why.
I think Google search has just declined a lot. I guess they’re losing the constant cat and mouse game with SEO. It seems worse than it has ever been, I’m relying more on ChatGPT and copilot now.
I can only imagine that LLMs will be the end of any content based search ranking. I don’t know how they’ll adapt to that.
This is why Google has basically surrendered and why so many search result categories are now dominated by whatever sites Google has arbitrarily declared the winner through editorial decision making. In many search categories, we are effectively back where we started with Yahoo directories and hand-picked search rankings. What you see on the first page SERP is the best that they can do under the circumstances based on the fundamentals of how search works.
The web was a fun idea while it lasted, but if you are using it as a primary information resource, you are wasting your time.
I for one welcome back our new library overlords.
Published by Pearson (January 14, 2022) © 2021
Richard G. Lyons1. Heavily penalize presence of ads. The content farms makes their money with lots of shit ads. This will wreck their business model.
2. Just manually block out heavily penalize content farms and boost known good sides like SO and MDN.
There's just a lot of fraud on the internet related to search and advertising manipulation, but it's under-policed in part because of the internationalized nature of it and because it is hard to bring fraud cases in the United States because of the particularized pleading standard. That should not stop the feds from bringing criminal cases, but generally the feds care about large dollar value frauds (as they probably should be) rather than on policing very large numbers of small dollar value frauds that have a major aggregate impact on the online economy. They like going after the guys who steal $100 million from deaf children with Lupus rather than doing 200 $500k fraud prosecutions.
Duck-Duck-Go returns the up-to-date documentation link 2nd, and the MDN result in 3, with W3Schools in 1. Bing returns actual content on the results page, describing exactly what you need to understand a TS Array.
Google have the incentive to push the poor sites, because they earn revenues from doing so. Bing and DDG don't have that incentive, and return much more relevant and useful links. That doesn't feel like a coincidence.
Failing at what though? Is it anything they care about, that they want to do?
If not, then it's not so much failure as it is a change of plans on their part. They don't want to do that anymore, and there's no one else to pick up the slack.
https://kagi.com/search?q=typescript+array&r=us&sh=Qa5cXHvwj...
We simply downrank sites which display a lots of ads on them and also use community blacklists for dev site clones.
Unfortunately it still doesn't solve the issue that sometimes the good results are still buried pages away or simply not come up at all due to google's shitty algorithm.
I really need to look into SearXNG or something.
This will hurt them long term I presume, but they won't care because they earned money.
Now, I do have a Stackoverflow user as well, but I actually prefer publishing my ideas on my own site rather than help build someone else's content farm for free. Stackoverflow is, itself, a content farm, and it can be very hard for new users to join the site. You can not even post comments without first earning enough points. For a very long time I would actually resist joining the site for that reason. I have only recently earned enough points to comment.
Now, I happen to own a so-called "content farm" too, and the choice can either be between creating a standard blog with very little traffic or try and cover everything you can possible think of in order to compete with other "content farms" in your niche. It is very difficult if not near impossible for a single individual to create a valuable resource and maintain it, and it is simply not sustainable if you have paid authors working on it as well. There is no way you can monetize it decently. Stackoverflow probably found a way around this problem by simply leaning back and monetizing their users content.
Once your site grows big enough, you also deal with a ton of spam- and hacking attempts. Everything combined just requires an inhumane amount of time to deal with.
Of course, authors are desperate because of how difficult it is, and perhaps especially authors from poor countries that might not have other sources of income. Their basic business model seem to be: create a content farm with ads, fill it with copy-written spam and hope Google indexes. Often these sites even have multiple authors, which is quite baffling given the extra expense it must create for them. But I do not think they have actually thought the idea through – because it is just not profitable.
Weirdly it's often in the technology niche, which they are clearly not proficient in, and more or less containing stolen solutions with little original content added.
I have seen a few sites like this, ripe with some of the most nasty grammar too. It interesting they are able to rank simply based on their volume? Of course they must be using blackhat techniques, including linkbuilding if you analyze their link-profiles, because there is no way that something so poorly designed and maintained is getting that much attention compared with official sources or stackoverflow.
For those of us who own blogs, such sites are often easily outranked simply by writing a comprehensive article on whatever tiny topic they have posted about.
I have no allegiance to SO ownership so when the fake SO sites show up in results instead of SO, usually reading them will just give me the answer more quickly than finding the actual SO source.
There's a good reason for that. Sites come and go and as a result links to solutions die and you wish someone had just answered the question instead of just linked to it.
now, in an ideal world competition would solve this problem, but the cosmetics companies heavily collude and anti-compete to prevent this
I suppose the real frustration is that Google became so pervasive that bookmarking a website and using its own search functionality is a total afterthought.
The problem is, they won't, because active moderation beyond responding to legal (DMCA, right to be forgotten, anti-CSAM) demands would massively endanger their "we are an impartial search engine" defense.
SO had an expert sexchange.
1. Respect exact match searches - this used to work enclosing the search terms in "" quotes, but no longer works. If there are no exact match results, return nothing.
2. Allow blacklisting or removing results from certain websites entirely e.g. I want to be able to configure geeks4geeks to never show up in any results ever
If someone could make this new search engine they would have a good shot at replacing Google :)
probably my new default search engine. Thnx
We are not trying to replace Google though, but offer an alternative to people who care so much for the quality of their search experience, that they are willing to pay for it.
[1] https://kagi.com
And if you find a result that got included despite not being an exact match you can report it and see it get fixed in a few days.
To be honest I think your pricing is to high. $25 for unlimited queries might be fine for somebody who needs a good search to work and earn appropriately.
But as a (former) PhD student I ran through the 100 free queries in 2 or 3 days and just would not have been able to afford 25€.
I would gladly pay 10€ (for unlimited searches) or 15€ (for a unlimited family option). But to me, 25€ just seems to high. That's 5 meals at my workplaces cantina right now (Germany, NRW).
(I assume you are aware of pricing issues as pricing options have changed at least once while kagi is on my radar)
Unlimited for $10 is something we are working towards.
Also thanks for creating kagi. Kagi was the first "alternative" search that convinced me that there can be competition to Google. YaCy just does not work, most competitors (DDG, etc) just repackage the big engines. I use presearch as my daily driver right now, but am somewhat turned down by the NFT shenanigans behind that. Kagi looks like the only engine that stands on its own and is definitely something worth paying for.
Perhaps 10 years ago Stack Overflow was able to do some minimal SEO and then get by on content strength. Perhaps nowadays Google is doing a good job preventing basic stuff from working, so the only people to get good results are SEO-ologists that only know about exploiting SEO, and have nothing interesting to say on any other topic.
The competing answers paradigm is also fundamentally broken, I don’t want to see 15 answers to “How do I get the length of a string in Python?” I just need to see
len(x)
Programming splogs do better than SO does in this respect. In fact, even the Q/A paradigm is bad because the average SO post requires scrolling past at least one extensive code example that does not work,For more than 10 years I thought the world needed a search engine for programmers. You really ought to be able to upload your POM file or equivalent and have the system automatically search the correct version of documentations. (Any attempt to look up things in the Java manual has to be written like “JDK17 javadoc {className}”; Javascript libraries like reactstrap, react-router and such often have a few wildly incompatible versions and I don’t want to waste a millisecond with the wrong version doc, …)
I wouldn’t mind searching answers from stackoverflow but I only want the best correct answer and I don’t want to read a long confused question, etc. As this would clearly save coders time maybe they’d pay for a subscription as they do for Jetbrains tools.
Maybe there's a point where the internet, with decades old information pilling up, becomes unbearably big for indexing services to handle all of it in a efficient manner. Hence the recent "optimizations" that companies swear haven't worsened searchability.
The flip side of that is a large proportion of those are no longer fine or operant.
I already go to ChatGPT to cut through the SEO-optimized crap that Google offers me in the first couple of result pages. I would bet that a lot of the responses given by ChatGPT come from Stack Overflow.
Now, what if we had StackGPT which offered me similar funtionality as ChatGPT, but better? E.g. respond with some code and an explanation, but also link to the sources (which are probably within their site, so they have prime access to them). Or offer as an explicit option to respond using sources other than their archive, but perhaps without citing sources.
Are they just fine today, too? To judge that, you have to look at the date of the question and its answers, make an educated guess at what OS/language/library versions they are about, judge whether that makes a difference for the version(s) you’re using, and only then evaluate whether the reply even was correct at the time (it may have had a thousand upvotes, but still be dated)
I think a really good Q/A resource would require posts to be tagged with version info. Most people think manual tagging isn’t fun, though, so it’s hard to get such a set from volunteers.
An alternative would be to require test cases that the site can run to check what version(s) replies are valid for, but writing such tests that do not break over time is hard, and, again, in general volunteers don’t like writing them.
That leaves generating tags or test cases. I don’t think we’re there, quality wise, to do that.
"how to load a file all at once in python" returns a first hit pointing to a blog post answering the question correctly, a second pointing to a SO answer that is actually for a slightly different problem but contains the correct answer, answer #3 is a youtube video that probably answers the question correctly.
Geeks4geeks doesn't show up until #4, well below Stack Overflow. (FWIW, their answer was fine too).
> You seriously can't find useful anything these days using Google.
That really feels more like a meme than reality. Are there other subareas where the SEO is doing better than this one? It seems like a pretty representative question.
Oh, so maybe DDG is ahead of Google in terms of quality alone in at least one domain now? Because it definitely still gives answers from SO at the top.
I haven't noticed that at all. I also haven't noticed any degradation in the results on Google overall contrarily to what people claim.
What I have noticed though is that every time someone complains about search results and shares the way they search on Google, it becomes obvious that the problem is between the keyboard and the chair and not in Google.
For example my girlfriend always get completely unrelated results when searching. But she searches an unintelligible mess that even a human wouldn't understand.
I’m not a share point developer, at all, it was my first time with it. I am of the old school however, so figuring out how systems work by reading the manual or specification isn’t foreign to me. I’ve worked with the bitmap format, I’ve worked with solar inverters, I’ve worked with embedded construction software and so on, so sharepoint should’ve been easy. But I couldn’t have done it without StackOverflow.
Fast forward to 2023 and I need to build something for sharepoint again. Haven’t touched it since I was sort of dreading it. Only this time the official documentation made it so easy I never needed anything else. I’m sure StackOverflow could have helped me, but I didn’t need it.
Those are very isolated examples, but really, that is my personal experience with almost every thing I work on these days. Yes, I’ve also got 10 more years of experience under my belt, but I do think it’s because we as an industry have become much better at working through the official channels. I mean, when you have a problem with something today, so you go on StackOverflow or do you go to the GitHub issues (or whatever else they have)?
I'm not 100% sure this is the fault of the API, if it's because of SharePoints indexation, if it's because of the 3rd party meta-data indexation that we buy for our document library or if it's because of how terrible our own architecture for the data flow is. But I've had to build a lot of redundancy and caching into the terrible piece of gaffa tape which is our integration, because SharePoint won't give me every document everytime that I ask for them.
So a bad experience? But at least it's not FTP (yes not sftp) pulling different file formats from 9000 solarplant inverters bad.
I never had problems finding SO answers for C++.
One thing I noticed in the first chart displaying visits is that there is a sharp drop in April, which is just after Google deployed its last core update: https://status.search.google.com/incidents/Cou8Tr74r7EXNthuE...
Q: [very explicit title about F in context X] I'm specifically not asking about F in context Y which I know about but is irrelevant; I'm specifically not asking about libT which may accept it in accordance with standard U; instead, in context X, is F working for libS?
A1 [100+] [+50 bounty] [auto-accepted] [date d]: "it works" + long winded answer about decontextualised substring of F proposition working for libS in context Y
Comment 1: this is not the correct answer, see A15
Comment 2: complaint that A15 note about U is from a second hand site so answer is warrant of discredit [even though authoritative docs about U are not publicly available]
A2 through A7 [30+]: "it works" (rehash of A1, posted within d+[1..365])
A8 [30+]: Q is a dupe, see this answer [link to question and answer about libS in context Y]
A9 through A14 [30+]: "it works" (gives example for libT, which behaves differently)
A14 through A17 [30+]: "it works" (gives example of standard U, which libS does not comply with)
A15 [5] [d+5]: Q being explicit about F for libS in context X and explicitly _not_ about context Y, let's address it for libS in context X: *it does not work* because foo (link to libS doc || source code || example). Tangent for the sake of completeness, context X is irrelevant because bar is orthogonal to foo + context Y. standard U does not apply because libS is not compliant [quote or link about U]
Comment 1: this is the correct answer
Comment 2 [d+1095]: answer is invalid because libS vN+1 has just been released and adds F in context X
A16+ [<0]: completely haphazard lottery answersI don't think SO "auto accepts" answers. The only one who can do that is the original asker, manually. See: https://stackoverflow.com/help/accepted-answer
On meta.so this has been asked repeatedly and the consensus is that no, it cannot be done and furthermore, "accepted" doesn't mean "the best", it just means whatever the original asker marked as helpful for their situation, so "auto-accepting" would be meaningless. Also see: https://meta.stackoverflow.com/a/262915/147346
As for upvotes: if a question is upvoted that means the community upvoted it. The "community" in SO is just regular people, programmers both knowledgeable and newbie. I don't know that any "open access" platform can solve the problem of people upvoting the "wrong" answers; it's not a tech problem. It happens here on HN, too!
Of course the users themselves could have deleted the questions, but regardless, I did not expect to see that.
It could be a nice experiment to enable RSS for some niche topics and checking automatically after a number of days to see which are gone.
Stackoverflow once set out - with the very best intentions - to be better than expertsexchange et. al. I think that's why they are so strict with their moderation.
The issue, of course, is that this is a hard problem. To moderate a community of experts you need to be an expert yourself. With Stackoverflow I feel most moderators are just overwhelmed with the task, so they aggressively shut down everything they don't understand.
Many of them are probably aware of that and just don't know how to handle it better. Others are so clueless that they do not even recognize the good questions.
Can't agree more. So many relevant questions shot down by an anal retentive mod you end up having to find the answer on another site.
Related: I've often been looking how to do X, find an SO question asking that, but the answerers there refused to answer until the person explained why they wanted to do X, and then all the answers (correctly) told the person that they actually needed to do Y and explained quite well how to do Y.
I actually need to do X, so those answers are useless to me.
Then I find another question on how to do X, and the mods close it as a duplicate of that earlier question. Even when the questioner specifically notes in their question that it is not a duplicate of that earlier question because they really need to do X the idiot moderators close it.
I'm thinking of things like assigning constants to pointers in C and/or manipulating pointers directly.
> I actually need to do X, so those answers are useless to me.
I know what you mean. Whenever I (rarely) ask a question on Stack Overflow, I always have to defensively load it up with language anticipating misinterpretations and instructing people to answer my question and not some other one.
Otherwise, internet-point-chasers will fall through the woodwork giving easy, worthless advice. Even with all the defensive language, a few always show up.
Agree. This isn't a problem on the partner/sister sites as much. They are very helpful too. But they don't show up on Google
The best you can hope for is some answer down the page that says something like "to answer the actual question..."
I call them pseudo-trolls because I think they are well-meaning, but they function as trolls: overrunning a web site, hijacking discussions with repetitive and irrelevant content, and making most potential users feel that participating isn't worth the time and effort of interacting with them.
So I can't add relevant info/feedback to a topic I landed on via google, because I don't have reputation, so in essence the site strangles itself. Because in 2023 I'm not going to farm(!) 50 points just to share my info. All this because a few bad apples tried to game a system and someone came up with this lame idiocy of you must have X reputation...
But you can ask questions.
But if noone upvotes your answer because it's a hard one, and/or noone knows the reply to your problem, then you are stuck. So should you ask some blatant question, like, is the sky blue and why? But it will be a duplicate. No reps for you. So it's a totally braindead system, they deserve to die.
Why don't they introduce a system where you can say, I need this answer for 5 bucks. But that will not be implemented, because it would make a dangerous precedent where you could harness other's valuable knowledge in almost an instant.
I think the problem with that is how do you prove it's a good or bad answer? Say you offer a bounty and I generate a great answer. You look at the answer and use it and then flag the answer as bad so you don't have to pay.
Yup. And that's why I stopped using SO quite a while (years) back.
That's not because I was intentionally avoiding it, but a natural side-effect of it becoming much less useful, largely because of the thing you're citing here. The quality of answers has fallen quite a lot, and SO seems to actively avoid hard questions (and if I'm searching a site like SO, it's because I have a hard question.)
SO burried itself
Some hot shot user with an absurd amount of badges told me, in a condescending tone, that my solution did not answer the question and linked me to some TOS article or something about staying on topic. This was the first time I had ever answered a question on SO.
My takeaway from the situation is that SO is full of accounts that farm badges/rep. To what end, I do not know - perhaps they reference it on their resume or portfolios.
I called the guy out, seems he has since deleted the comment.
Sure, you may encounter the occasional rough moderator but having to correspond with random people on the internet in writing as my daily job, I can't blame them. If you don't have a rep, it's up to you to gain some by crafting the best question, linking to the relevant docs, transcribing screenshots, correcting typos etc.
The only issue with SO imho is that it's getting too big and there should be a lot to gain from further splitting to other StackExchanges like computer science, computer graphics, databases, infosec, Vim etc.
It would be great if there was a way to transfer a SO question over the more relevant community I guess
It was irritating as hell to have the top several Google links go to SO discussions that had been shut down by the moderators.
> I can't blame them
I'm sorry, this is like saying you can't blame cops for beating up the occasional suspect because they encounter a lot of genuinely bad people every day.
If you can't do the job in the presence of shitheads, you have no business being a cop. Or a forum moderator.
There are lots of ways to respond to a positive indicator that could reflect a aelf-awareness of your own false positive rate. Being trigger happy probably isn't the correct response.
It isn't clear how SO moderates the moderators, do they have an IA team that investigates corruption?
At least for "closed as duplicate" that is working as intended: "Closed" doesn't necessarily mean "you shouldn't have asked that question". In the ideal world, a post being closed as duplicate just allows people that phrase the duplicated question differently to also find it.
Also Stack Overflow is not supposed to be a public service unlike police (is supposed to be), so the moderators are not accountable to you, only to the site guidelines (which you appeal to I guess)
Given the value it gives to society, it should be turned into a public service before some rich prick takes it over and turns into a cow to be milked or a self-aggrandizing venture.
Read again. It has nothing to do with "ego". By allowing their "moderated" threads to continue polluting search engine search results, they are wasting my time. They are wasting the time of everyone else who has a similar question and makes the mistake of clicking through an SO link.
I would not be at all surprised if they got downranked because people at Google were tired of SO wasting their time.
The rough moderator problem exists on other SE sites also (source: personal experience) where they are IMO needlessly rude, snarky or dismissive.
Especially nice to be able to link my own answer to things.
If you are either someone who likes wielding power recklessly or autistic then becoming a power user isn't a problem. But demanding users "prove" they are trustworthy in a system where blatant abuses of authority go unchallenged is farcical.
The wheels of this kind of stuff turn slowly, but obviously ChatGPT was trained partly with data only gainable from a healthy StackOverflow-kind of site with users actively asking unique questions and enough people answering those unique questions with well-though-out answers. The shittyfuture outcome is that StackOverflow goes out of business and LLM's stagnate on this front, while being capable of answering in fuwwy speawk when prwompted, would still be limited to / biased towards older versions of libraries, software, languages, tools etc.
Claim asserted without evidence.
It's painfully slow. I can Google the question, click one of the top results, skip to the relevant part and read it faster than GPT can generate two sentences. You also have to build an elaborate prompt instead of throwing two/three keywords into it.
It doesn't help that GPT is insistent on replying in the three paragraph format, meaning that the first 30-40 words it creates are just trash to be ignored.
I found it useful once - when I had to write an essay about ISO 27001 for college and just wanted it to go away. Took what it generated and spent 20 minutes editing it to look closer to my style. For real work it isn't as useful.
I don’t even open SO anymore; if it has a direct answer to your question, the LLM almost certainly does too; and asking new questions on SO is basically impossible.
If you manage to survive the gauntlet of “too specific, already answered, not general interest, arbitrary moderator activity”, the chances of getting an answer that answers your question can take forever; most likely you’ll get a stupid answer that doesn’t answer it, upvoted by idiots who don’t understand that it not an answer the the actual question, and, ultimately, because it “already has an answer”, ignored, never to receive an answer.
Maybe one day, a passing savant will answer in a comment.
…and yet, you find it faster and more reliable?
You, and I, have had different experiences on stack overflow in the last two years.
> For real work it isn't as useful.
To you.
Which was the whole point of this chain :)
Ironically, this is why people like me prefer LLMs (when they're accurate). With Google, about 50% of the times the top SO hit is not answering my question. So I have to click 5-10 SO links, parse each one to see if:
1. The question being asked is relevant to my problem.
2. The answer actually answers it.
I may be able to do it quickly, but it is a tedious burden on my brain. While GPT doesn't always work, the nice thing about it is that when it does work, it has taken care of this burden for me.
Also, GPT's pretty much memorized a lot of the answers. I once asked it an obscure question involving openpyxl. It gave a perfectly working answer. I wondered: Did it reason it and generate the code, or is there a SO post with the exact same answer? So I Googled it, and sure enough, there was an SO question with the same code!
Except GPT's solution was superior in one tiny respect: The SO answer had some profanity in the code (in a commented line). GPT removed the profanity :-)
When I search for something, I want factual and correct information, which is not what a language model is for.
However GPT4 is occasionally still inadequate so I certainly hope SO sticks around, it is a tremendous resource.
I currently pay for it but I'm not sure if it makes any difference.
Ask it how to use it
My experience - the free version made up a lot of things but still felt very useful - enough to want to upgrade to the paid version. With the paid version, I notice very rarely that it hallucinates. It does make errors but it can correct them when I provide feedback. It is possible that I just do not notice the errors you would notice, it is also possible that we use it differently. I would like to know.
Paid.
> Would you kindly provide links to a few examples of very fundamental things incorrect?
No, definitely not.
> I notice very rarely that it hallucinates.
Unsure of what "hallucinates" means in this case. Some examples of things I've used it for: docker configuration, small blocks of code, generating a cover letter, proofreading a document, YAML validation, questions about various software SDKs. The outcome is usually somewhere on the spectrum of "not even close/not even valid output" to "kind of close but not close enough to warrant a paid service". When I ask for a simple paragraph and I get a response that isn't grammatically correct/doesn't include punctuation, I'm not sure what I'm paying for.
The term "hallucinations" is now commonly used for instances of AI making stuff up - like when I asked ChatGPT (before I had paid account) to recommend 5 books about a certain topic and two of the recommended books looked totally plausible, but when I tried to find them, I discovered there are no such books. This is where I see a big difference between GPT-3.5 and GPT-4.
>> I get a response that isn't grammatically correct/doesn't include punctuation
What punctuation? If you mean stuff like commas separating complex sentences, my English is definitely not good enough to spot that. But your mention of punctuation reminded me of problems that ChatGPT has with my native language... any chance you are using ChatGPT in a language other than English?
Here is an example of using it to write simple powershell scripts:
* https://chat.openai.com/share/17a9d185-4a8c-4c10-97a2-0c5b08...
Here is an example of using it to merge different types of json configuration files:
* https://chat.openai.com/share/ed0c9cf9-7586-43df-9271-7b5859...
You need to pay for it if you want access to the latest version of the model, along with some beta features like plugins. Plugins are extremely useful and it is worth paying just to get access to them. For instance, you need to have a certain plugin to get it to read links.
That's not very convenient, atleast for now. I got so used to search engines by now, that it only takes a few keywords to get the expected result. Be it a SO-answer or a documentation page. And as people have mentioned, chatGPT was learned on the stuff that's on the internet, so if there will never be any new stuff, because people just use AI, then it will not learn and won't answer your new questions. For some edge cases I might try AI here or there, but usually it's not for me.
Hell, there comes even an example to my mind. I recently just asked chatGPT what a single-issue 5 stage pipeline on a CPU actually means. I wanted to know if, especially, the "single-issue" meant that only one instruction is present in the pipeline at a time, or if a new one gets shifted in on every clock cycle (if there is no hazard). It just couldn't answer it straight-forward. It was also kinda hard to find the exact definition on the internet. I found it in a book from the 90s which was chilling in my book shelf (Computer architecture and parallel processing by Kai Hwang). Hint: Single-issue just means that only one instruction can be in one stage at a time, but still multiple get processed inside the pipeline. The keyword is 'underpipelined'
I'll just keep an eye on AI progress, but will probably not make it my goto for some time. Maybe later (whenever that is)
That's because you are comparing asking ChatGPT to write full code to searching for a question on Stack Overflow and adapting their answer (which is comparing apples and oranges).
Try using ChatGPT like you use Stack Overflow instead (i.e. the question is "How would I record an audio stream to disk in Python" rather than "write me an application / function which...").
As an aside, try "How would I record an audio stream to disk in Python"" in both GPT4 and searching for an answer on Stack Overflow and see what has the better answer! (Clue: GPT4, and if you don't like GPT4's answer just ask it to clarify/change it)
That's my point though. I get, that it can produce quite good results, if you are specific enough. And for some applications it makes sense to take your time and describe that as much as possible.
Most of the time I just need some small snippet though and usually I can get that with just a few keywords in my favorite search engine, which is way faster. So the conclusion is: There is no one or the other. They should be used complementary, or atleast that's what I am doing (as in use the search engine for quick hints and chatGPT for some more verbose stuff 'write me a parser for this csv in awk'.)
Plus I can ask follow-up questions in a context-driven way ("Can I do this without importing a library?").
I'm aware that different people will have different feelings on this though and personal tastes will differ, but while search engines stagnate I suspect the needle will continue to shift towards AI.
I can ask for recommendations for tools and libraries, which IIRC SO disallows.
I also don't have to pray my question will get enough vote attention or worry that I posted it at the wrong time of day.
On the whole, going the GPT route has been more satisfying in all ways.
Bing Chat almost always is useless for me with these kinds of queries. A few days ago I asked for a tool that monitors to see if a website is up. I told it I needed the tool to be something I'd run locally - not an online service and not something I need to sign up for.
It gave me 3-4 online services.
I reminded it about my constraint.
It gave me 3-4 more online services.
I reminded it about my constraint.
It said it didn't want to talk to me any more.
Stack Overflow is the programmer's internet bloodsport.
Better than reddit where you get insulted and banned by an actual LLM.
How is this being enforced? It's either bots banning bots in a digital game of whack-a-mole; or humans arbitrarily trying to asses whether something has been written by an LLM or a human.
There are some subjective signs that a post is LLM generated, like being overly verbose and making unrelated assumptions, or mix of horrible and perfect grammar. Those bans are hard to justify because the false positive rate is high.
But other signs are pretty obvious. My favorite is the use of APIs that should exist but don't. Passing parameters that neatly solve the problem but have never been accepted, or importing non existent libraries. I'm happy to flag those.
See: https://meta.stackexchange.com/questions/389811/moderation-s...
And I think Stack Exchange needs a new CEO. Maybe new owners, which is since 2021 Prosus. My impression is that they don’t understand what is the purpose for developers.
I can think of a few theories that I don't think hold water:
1. The rise of ChatGPT to answer many questions that StackOverflow would previously have been used for. This seems unlikely, since the timing doesn't really work out.
2. The perennial complaints about StackOverflow's culture of closing everything as duplicates or offtopic. This seems unlikely as well, since those complaints have been common for a decade or more.
3. The prevalence of SEO-optimized scrape sites - the ones that pop up with a "blog post" merely reposting a stack overflow question + answer in a different font. I've seen these for a while, and anecdotally they feel more common that they used to, but I couldn't give any real timeline for that vague feeling.
4. StackOverflow internal politics? I've seen the occasional stack-overflow meta thread pop up periodically on HN or social media, but I don't recall anything earth-shattering recently.
5. Most questions have good answers now and there's less need for new ones. I'd have bought that answer 10 or so years ago when StackOverflow's pile of questions + answers reached maturity. I don't think it suddenly hit some sort of answer saturation point in 2021.
My guess is that it's a slow shift in the culture of the StackOverflow userbase:
- Being a top answerer confers some cachet and makes you more employable in some places
- People notice this and start looking for the most effective ways to become a top answerer
- The most effective way is fast, low-effort answers
- There's been a rise in such low-effort answers over the last 5 years or so
- As a result, the cachet of being an StackOverflow top answerer is a lot lower
- The really good, deep, technical answerers (as well as the mods) are leaving as that cachet goes away
- Post quality starts dropping around 2021 and views start declining as people react to that in 2022.
I would choose first result - people are lazy. The problem comes and goes but it is true, spam sites, with content copied from SO, are ranking higher than SO itself. And it is often hard to tell you are viewing a copy on first glance, the layouts are usually more like a forum.
https://techcrunch.com/2023/07/20/tech-industry-layoffs-2023...
If I were building a Q&A site that genuinely wanted to encourage high-quality answers, I'd implement something like a 1-hour window where all answers are invisible immediately after the question is posted. This is to give people time to work on a good quality answer, without racing to be the first and gain those precious early upvotes. The UI could still indicate how many other people have answered / are answering, and if answer volume became a problem you could perhaps block new answers during that window after the first 20 or so have landed. When the hour is over, answers are displayed in random order and the existing site mechanics around upvotes etc kick in.
It feels like that could potentially address the problem of low-effort answers killing off the good ones.
Not to mention how beneficial it is for brands and products to create such closely-knit communities for their users. There's no need to Google things or ask on SO, as one can chat in real time and go back and forth for solutions with other users.
People vote to close and then move on. I don't mind editing my questions to satisfy the moderator's demands, but not if it has zero effect and just wastes my time.
I think people should be required to confirm their close/down-votes after a question or answer was edited, or else the votes should be automatically reverted. This mechanism should probably only apply within a limited time period after the question/answer was posted (or until a certain period of inactivity has passed).
It saddens me deeply, but currently I prefer not to ask questions, because the experience is so jarring.
- do not post a "teach the man how to fish" answer
- closing this question as duplicate because i'm in a hurry and i didn't read that the poor poster has already mentioned the possible duplicates and explained why theirs is different
- do not post an answer with links to the original docs that have context, instead copy/paste 3 lines that don't explain enough here
Basically... it's the moderation.
The other aspect that grinds my gears is the closing of duplicate question. Fine in principle, except when the original was answered 10 years ago and all of the answers are jQuery.
Later on when I had more experience, I wanted to give back:
- try to answer -> "you need cred"
- okay, upvote -> "you need cred"
- try to comment -> "you need cred"
- Ask a question -> "Duplicate. You should RT(Free)M we built."
- Last month, Stack Overflow had ~142,575,642 visits.
- Last month, Google gave 127,896,508 visits to Stack Overflow
- Last month, Bing gave 7,491,274 visits to Stack Overflow
You could say that the Stack Overflow # of pageviews depends mainly on:
- How often people are searching Google/Bing for answer.
- How often Google/Bing rank Stack Overflow high enough for people to click into it.
https://medium.com/snowflake/how-to-load-the-stack-overflow-...
The Hugginface repo unfortunately prefilters some of the tables/rows according to some criteria, making it less usable for general analytical queries that the BQ or SEDE datasets enable. If anyone knows of an 'XML-streaming' solution that directly samples from the Internet Archive's data dumps, I am all ears.
[1]: https://huggingface.co/docs/datasets-server/rows
[2]: https://huggingface.co/datasets/HuggingFaceGECLM/StackExchan...
For instance, this user query by Starball tracks network contributions over time[2][3].
[1]: https://data.stackexchange.com/meta.stackexchange/queries
[2]: https://data.stackexchange.com/meta.stackexchange/query/1759...
[3]: Static image if the query times out: https://i.stack.imgur.com/LYZQm.png
What non-trivial question has straightforward answers? It's just a homework help site. No serious discussion is allowed.
Why would anyone other than undergraduate students spend their time on such a site
Every time I started out learning a new framework, SO would be tremendously helpful, because in software development, you mostly have questions that have correct and incorrect answers.
Maybe software engineering is all solved and there's no more ambiguity left to be discussed.
I'm sure everyone involved actually wants the site to be useful and pleasant, but somehow the actual result is the exact opposite.
That being said, I don't think the culture or moderation has got worse recently, so I suspect the traffic decline is either a change in Google's algorithms or the impact of ChatGPT (or both).
Once I was researching something and found wrong answer on SO from some junior dev from India. The answer was wrong and it was evident that the author of the answer didn't even fully read the question. Out of curiosity I checked the author's profile and found out that it was a deliberate strategy: e.g. when there's a question about working with files, he'd post an answer with a link to the language page on "open".
I politely asked the author in the comments to stop posting wrong or useless answers. Instead, he wrote very aggressive response, opened my profile and downvoted as many of my answers as he could.
And there were many similar cases, probably still are.
Which is aligned with the site owners' point of view, so they have no incentive to change anything. Though everyone notices that quality of articles and comments started to decline once the users with karma majority censored all others out.
These are constraints that limit the types of people that are interested in remaining here. But HN is by no means "small".
I assume that there will be a decline in quality. It will become an interesting case study.
So in short, gambling.
And the house always wins.
Let's not forget that in india it is not considered wrong to "game" the system, even it is considered unethical by western standards. It is not a coincidence all of these tech support scammers are based in india.
this brings forth an interesting topic, how do we reconcile cultural differences in online platforms ? Nobody can deny they exist, and they create clashes like the aforementioned one.
When it becomes apparent that there are cultural norms that make it difficult or impossible for people from those cultures to use your platform, you re-examine your code of conduct, of course, to make sure you're not unduly excluding people. If you are, you amend the code, you change the platform in ways necessary to support them, and everyone tends to end up better off as a result.
But some cultural norms will be antithetical to the point of the platform, like in this case: SO cannot function properly when people attempt to game the system. In those cases, the CoC remains, and people from that culture will either have to adapt to the platform or find (or make) a similar one that works for them.
The alternative, as we see here, is that your platform degrades until it's not useful to _anyone,_ this "problem culture" included.
gaming the system is prevalent but not considered ethical in india. HRs won’t allow such candidates for example. So there is no question of amending CoC to make room.
There is a difference between something being prevalent and something being culturally accepted. The solution lies in naming and shaming. Making it clear that such people aren’t wanted on these platforms.
But i don’t think anyone should chalk this up to “cultural norms” That’s just being unfair to those of us who are just as fed up with this mindset.
Do people not remember the "Hacktoberfest" fiasco? People were encouraged to make a PR on public repos in exchange for a free t-shirt, and _obviously_ this led to repos being overrun with spam PRs changing a single line comment.
https://www.theregister.com/2020/10/01/digitalocean_hacktobe...
There were dozens of videos like this one,
https://www.youtube.com/watch?v=8xjmCsdgUhE
with loads of people replying and thanking the author, excited for their t-shirt, not really knowing they're _creating unnecessary work_ for actual engineers with limited time.
Incentivized systems like this can work okay in small, self-selected communities, but as soon as it's open up to the entire world, the implied courtesy of said system vanishes, and immediately we're pandering to the lowest common denominator.
This thread digs in a bit. https://twitter.com/Kautukkundan/status/1311717814768594944/
I've a rather high account. So I can revert mod decisions and sometimes do. E.g. reopen a marked-duplicate that isn't duplicate on closer inspection.
I don't do it often (can't be bothered too much) but when I do, I sometimes almost feel the hate radiation over the web. Indeed, people on a rage-tantrum going through all my answers and questions and downvoting those. Lol.
But SO (the company) wanted it to be more accessible, easier for newbies, "nicer", there was a huge uproar over them publicly blasting a moderator over a disagreement on a unilaterally imposed new code of conduct, and recently they even (again unilaterally) effectively reverted the ban on LLM-generated content. This has been going on for years, and moderators have less power than they ever had. Imho this whole thing started much earlier, I think it was 2017 when they tried the SO documentation project and let everyone keep their rep where I first thought they had jumped the shark.
The company has been on a bender for the last few years, and high-ish rep users like me just don't see the value in answering anymore. With recent blog posts it seems even more clear that they are on the direct path to enshittification, against their own actual users and moderators.
Recently it feels like all the actual professionals have left and what remains are "students" asking low effort questions in bad faith while an army of spambots tries to pounce on these.
While this means boosted engagement numbers short-term, it spells death of the site long-term. SO cannot survive on low quality spam, even if other sites (like Reddit) may be able to.
There's a question on meta from a decade ago about what the Roomba (automatic deletion scripts) should be deleting. One of the answers had a bit of point in time information:
https://meta.stackoverflow.com/q/262077
> How much traffic do the questions that get duped to something bring? Especially the (currently) 410 questions linked to the Java NPE question.
The Java NPE question is https://stackoverflow.com/q/218384
It now has 10,339 linked to it. https://stackoverflow.com/questions/linked/218384?lq=1
That makes it harder to find the original source, more difficult for google to chase, and clusters results.
> The original idea is great, but like any other place it needed to evolve barriers as userbase skyrockets and it mostly failed to do that so far.
Which upper management has been working against by encouraging a reduction of those barriers, disincentivizing people from moderating and curating, and trying to grow engagement from people asking questions.
When it was small, it was easier to handle the questions of the day, guide new users, handle the "fun" questions (and answers) in a way that wasn't off-putting, and generally be a smaller community.
As SO grew, it lost control of the culture that had been established there before (much like Usenet (different thread)) and became a place for people to do hit and run questions - drop their question, come back later, get the answer and move on.
The majority of the users of the site had moved from "community of people sharing information - asking and answering" to "new users without any cultural attachment asking a question and not remaining."
The core group culture became more defensive of their ideals... and lots of friction between management wanting more engagement and new users just wanting people to answer their question ("if you don't like the question, just move on" being a frequent refrain).
From A Group is its own worst enemy:
> 2.) The second thing you have to accept: Members are different than users. A pattern will arise in which there is some group of users that cares more than average about the integrity and success of the group as a whole. And that becomes your core group, Art Kleiner's phrase for "the group within the group that matters most."
> The core group on Communitree was undifferentiated from the group of random users that came in. They were separate in their own minds, because they knew what they wanted to do, but they couldn't defend themselves against the other users. But in all successful online communities that I've looked at, a core group arises that cares about and gardens effectively. Gardens the environment, to keep it growing, to keep it healthy.
> Now, the software does not always allow the core group to express itself, which is why I say you have to accept this. Because if the software doesn't allow the core group to express itself, it will invent new ways of doing so.
As the software didn't allow sufficient and proper moderation and curation tooling, the way that the the core group expressed itself was the more negative and ultimately toxic approaches. Snark and rudeness are the moderation tools of last resort.
A Group continues with:
> 3.) The third thing you need to accept: The core group has rights that trump individual rights in some situations. This pulls against the libertarian view that's quite common on the network, and it absolutely pulls against the one person/one vote notion. But you can see examples of how bad an idea voting is when citizenship is the same as ability to log in.
The current goals of upper management being advertising and engagement are not in alignment with the goals of the original founders (as idealistic as they were) and what remains of the core culture.
https://blog.codinghorror.com/introducing-stackoverflow-com/
> It is by programmers, for programmers, with the ultimate intent of collectively increasing the sum total of good programming knowledge in the world. No matter what programming language you use, or what operating system you call home. Better programming is our goal.
Note that good is italicized in the above quote and is present in the original.
---
Getting people to be able to find existing questions and clean up the SNR of the content out there on SO would improve it... but that would likely make a lot of lines go down rather than up (deleting 10,000 duplicates of one question would show up).
Trying to get Stack Overflow back to a scalable model doesn't further the engagement and upper management goals directly.
Instead, they're focused on more engagement... with not unexpected responses.
https://meta.stackoverflow.com/q/425532
https://meta.stackoverflow.com/q/425531
https://meta.stackoverflow.com/q/425530The first is marking so many interesting questions and answers as off-topic or opinion-based, which vastly reduced the scope of interesting content, and discouraged daily visiting to learn instead of just using the site as a reference through google. Relying on google to feed you users is a fool's gambit. Websites need to create recurring visitors, and they do that by providing something that engages. Stackoverflow and software.SE's moderation policies made it far less engaging as they grew more heavy-handed. This left the sites open to changes in google's algorithm affecting their traffic.
The second major problem is duplicate policing and a lack of staleness policies. By not allowing the same question to be asked again it meant that the content of the network has grown gradually more and more stale. Even my own old answers are now often wrong because they are simply outdated. Stackoverflow is filled with questions in the style "what is the right way to use technology X to do Y", and the right answer to that changes every 2 or 3 years, when technology X gets an update or when new insights in how to best do Y are formed. I tried updating a few here and there, when I saw people engaging with wrong answers, but overall the blame is on stackoverflow and its moderators for not working out a more effective mechanism to get rid of stale answers and letting new users answer old questions with answers that have a shot at rising to the top. This also means that as top-voted answers grow more and more stale there is a tipping point where google no longer sees the site as a useful resource and starts lowering it in the ranking, and this is what seems has happened.
It is still the best place to find a lot of answers (though their piece of the pie is rapidly shrinking), but participating in the system was never fun.
What accounts for this?
Although there appears to be a Monica effect (2020) in the data too. I stopped contributing myself around that time:
https://meta.stackexchange.com/questions/340906/update-an-ag...
It is also possible there's a methodology issue. There's no source for where this data comes from.
I scrolled down a lot and didn't see any scraper site links.
I can often evaluate a Copilot autocompletion (check that code looks right at first glance, check that it compiles, hover over method+type signatures to see their docs, run the code) in less time than it would take me to find+read a Stack Overflow answer.
Code is easy to verify anyway.
IMHO people have tired of the smug dismissive assholes that dominate SO and its its insipid gamification and they have, gradually, found alternatives.
For some, it's more welcoming and knowledgeable communities in github issues+discussion.
For others it's special-purpose, interactive forums that provide more guidance and non-hostile support for users of certain platforms/tools (eg Posit Community).
For many, copilot has been fulfilling that need, Does it always give "the correct" answer? No. Does it provide a sketch of a solution that gets you half-way there when dealing with tedious humdrum stuff that you just forgot because it's so boring? YES.
For yet others it's just plain old reddit where you can ask a question and whether it gets smacked down or not depends on how cool the community happens to be.
Until seeing this article, though, I hadn't thought about how much I had unconsciously changed my habits to work around the aforementioned issues -- I swapped to DuckDuckGo, started appending site:stackoverflow.com to searches, performing multiple searches around the same content (like grabbing a few keywords from a scraped article and then feeding that into a search of the official documents), and also more aggressively searching GitHub issues
With the caveat that I still use SO relatively frequently -- everything from 'I can't remember this specific syntax' to 'I have an intractable terraform/AWS/k8s issue and need to see if anyone else is experiencing it'
What was frustrating is that day in and day out, similar questions were asked - ones that are googleable.
One day I received a comment from such "user", under a question I answered and explained in detail. The "user" went to discredit me, without arguments or links and was pretty rude through all of the misspellings he did to avoid some automatic triggers. I replied rudely as well, asking him to read other people's answers before discrediting and providing no info.
Result was that I was sent to "cool off" for 2 weeks, and was given a speech from mod as if I was a child. At that point SO changed their ToS and claimed that all content was theirs, including answers I provided.
And that's where I departed from SO because from my experience - I was free debugging service for various agencies who went a little bit overboard with claims of their experience and in case someone was rude to me - I get no protection, however if I'm rude to "users" then I get slap on the wrist.
Working for free, being patronized, interacting with people who are too lazy to read - that were my reasons for abandoning SO.
Besides, when it started - all of the big questions about nearly every language were covered fast and there was very little to do except write SQL for other people and get praise for it as form of payment.
[1] https://aviation.stackexchange.com/questions/19569/how-many-...
[2] https://physics.stackexchange.com/questions/191826/derivativ...
Often itscthe same for css, where grid and flexbox are now standards, the old questions with replies full of vendor prefixes and very hacky solutions should be deprecated or transferred to archives.
Btw, how do the blog posts copying so questuons monetize and how does google not shut them down, the pages dont even work, I thought google would know better?
Another reason I rarely ask is of course as others have pointed out: the amount of pedantry is around 10x the amount needed to keep the quality of the site high, and often enough to make the quality worse. Often these days I ask a question, and get attacked despite it not being a dupe or off topic or low quality (I’m a top 1% account as I assume many of us are so I know my way around). Then a perfect answer comes and I can’t accept it because it’s already jumped on by mods. Somehow a useful exchange with a reasonable question and accepted answer has been completely drowned by mods. It’s infuriating.
At some point regular replies were replaced by administrators and mods who are seemingly far more interested in finding faults with your question rather than to actually answer it. The worst one was probably those times I’d ask a question, and then it got labelled “not a question.” Honestly, this unhelpful, bureaucratic, and down right nasty attitude has really disgusted me with the site.
Apparently new users are having even bigger problems, and they get flamed really hard for asking newbie questions. Often people are rudely asked to read the manual, when the entire cause of their problems is that the manual is so poorly written that it’s impossible to make any sense for it—even if answering questions like that is the “raison d’être” of a place like Stack Overflow...
Now, with the prevalence of ChatGPT and services that will give you great answers to almost any code-related question, I honestly think that the fall of Stack Overflow is well deserved.
Anyway, it's simply not true that "all questions are answered." Often when I've read the answer for a purportedly same question, the nuances of the other question is often nowhere near the original one, leading you to never fully having your question answered, which is extremely frustrating.
Often what you're looking for is simply a new way to explain an old concept, so that you better learn the inner workings of the problem. People are different, and so different people need different explanations. This is basic pedagogy. And that's also why RTFM does not work for all people. To think so is so arrogant that I'm at a loss for words.
My first question turned into a well thumbed tumbleweed.
My question got closed, with bullshit reasons. I kept asking to reopen over a few months, just to see what would happen (because clearly it was a good question and a good answer). After 3 months it got voted to be reopen. Since then I got badges for "popular question" and many upvotes there.
Pretty clearly I was closed by people who did not understand the question and probably had not even realized that I had answered it myself.
https://stackoverflow.com/questions/57064879/finding-coordin...
https://stackoverflow.com/questions/57323981/realtime-video-...
https://stackoverflow.com/questions/65025858/how-to-fetch-th...
Is it inevitable that traffic would decrease as questions are answered? Has traffic gone somewhere else? Has AI replaced it?
Now, AI kinda replaced for me, it's copilot first, then ChatGPT if copilot is not enough, then google/bing if ChatGPT is not enough and maybe finally SO.
Conservative figure from my side would be 15-25% less stackoverflow in a regular day.
I'm sure I'm not the only one.
Now I'm sure that ChatGPT probably scraped StackOverflow, perhaps using the very same answers I got from StackOverflow, but combined with other answers from other sources resulting in the instant answer. That does not necessary mean AI will replace StackOverflow, it just means that people wouldn't ask redundant questions in StackOverflow anymore, just questions that AI can't solve.
[0] https://stackoverflow.com/questions/477816/which-json-conten...
But I have asked a few questions, and the quality of answers have really declined, mainly because people are rushing to answer and not reading the question. They could address this by delaying voting on answers.
I'm sceptic that having having to choose between seeing upvotes or downvotes is helpful for the reader to draw conclusions. It would be nice to see this on the same chart.
It is hard to see how much voting and posting follows overall traffic trend.
But, posting and accepted voting had a huge uptick in May 2022, seemingly without any change in traffic. That is interesting. What happened?
In the overall traffic chart, it would be interesting to post some landmark events. Covid shutdown in US (assuming traffic is mostly US), release of ChatGPT 3
I'm not sure what the histograms in the tables are supposed to tell us.
Comment: "You shouldn't do A, it's better to do B"
Closed, duplicate of "How to do C".
Yeah sure, in the beginning, there were many more basic questions (how to increase number by 1, how to get division remainder, how to check if file exists,...), and if you're coding in Perl, you can still find all the answers... if you're working in python, you'll find answers for python2, some for python3, some specific to python2.6, etc... if you ask again, it'll be closed as a duplicate.
I know it's anecdotal, but after a few bad experiences, people just decide not to use that specific site anymore.
That makes it much, much less useful than it otherwise could be.
I would call that "stable" :) Or "good enough, that it doesn't need constant fixing" :)
My experience was that there was a long period where your complaint was quite true.
I believe that they have heard this complaint enough and have lightened up a little. There are still more and stricter rules than most sites, but the obsession with pruning duplicates seems to have cooled.
Stupid people don't realize they are stupid. When they can't understand a question they conclude there's something wrong with the question, not with them. And since they believe the question is silly, they happily close-vote the question.
This way they close questions which require specialized knowledge, experience, or deep understanding of the subject. Precisely the kind of questions I consider interesting to answer.
Sometimes I managed to undelete such questions, but SO made it easy to delete and hard to undelete.
Examples of the questions where I succeeded: https://stackoverflow.com/posts/57064879/revisions https://stackoverflow.com/posts/57323981/revisions https://stackoverflow.com/posts/65025858/revisions
In many more cases I haven't bothered. Sometimes, moderators killed questions so fast that I discovered that while trying to post my answer.
You probably assumed I asked that question? If so, the assumption was incorrect, I answered it.
> without explaining what you attempted, why it worked/didn't work, where you got stuck
I don't think any of that is possible to do. The question, and my answer, are too simple to decompose into smaller parts. Not enough runway to start and get stuck.
> expect other people to answer it fully
My expectation was rather different. I wrote my answer on Jul 16, 2019, and expected it to stay there.
Instead, on the next day some people have decided the question was bad, and closed it. Then on the next month, some other people have deleted it, along with my answer.
To close it, 5 people clicked once/each. To undelete it, I spent quite a few hours. Sadly, that's not a rare exception: moderation on SO is horrible.
The question is asking about the best way of doing X, which has a clear substep of "any way of doing X". Not only that, it's asking about a code solution, not just maths/combinatronics. Not a single line of code in a question about finding the optimal (not any) solution for doing some algebra in code is, IMHO, reason enough to close it.
Edit: for example, I'm not asking and so have no stakes on the question but can already try to think about brute-forcing it, which already shows more effort than the author SHOWS at attempting a solution.
To be fair, English is not my native language. But when I see a question “what’s the best way of doing X”, when it doesn’t have any criteria for what’s the best would be, and no other ways of doing X in the question, I consider “what’s the best way” part a redundant figure of speech. I view such questions an equivalent of “what’s a good enough way of doing X”.
> it's asking about a code solution, not just maths/combinatronics
Please read this: https://stackoverflow.com/help/on-topic According to that article, questions about math which don’t imply a code solution are offtopic on stackoverflow.com. They should be closed, and possibly moved to other stachexchange sites. According to that article, the OP’s question is good. The question was about a specific programming problem, and is a practical, answerable problem unique to software development.
> can already try to think about brute-forcing it
I’m not sure that’s actually possible to do.
- Does not show any attempt or willingness to try to solve the problem first on their own.
- Does not even give any indication of where the problem comes from, why it might be interesting, etc., it's just a "how to calculate X?", which could easily be a homework problem.
- It is about finding a (possibly) mathematical solution, and then implement it in C++. Two very big and different problem, asking the audience to do them both. Again no attempt to fix either of these two problems on their own.
- I'll concede the optimal thing might be a language issue.
> I’m not sure that’s actually possible to do.
But that's not my point, my point is that I already showed more willingness to try to solve this problem than OP. And THAT is a big problem. It's not on topic about any of those points, in fact if you remove the bit where OP is asking us to give them the full solution in C++ it could be a good question for the Mathematics SE!
None of that is required to ask questions on stackoverflow. For details, read “How do I ask a good question?” and “What types of questions should I avoid asking?” help articles. You’re inventing arbitrary restrictions.
Another thing is, “why it might be interesting” is subjective. Personally, I found the question interesting, that’s why I have answered it. You probably think otherwise, but note it only takes 3-5 votes to kill the question. Any question at all is guaranteed to have at least 3-5 people on that site who find it uninteresting, opinion-based, need more focus, duplicate, etc.
> Two very big and different problem, asking the audience to do them both.
Two big problems don’t have solutions which can be both explained in 3 short sentences. As you can see from my answer, the problem formulated in that question has such solution.
> I already showed more willingness to try to solve this problem than OP
You have not. However, you have demonstrated willingness to delete interesting questions based on arbitrary and subjective criteria, despite the question is perfectly in line with the stackoverflow guidelines. Which BTW is very on-topic, because I think that’s the main reason for the fall of SO being discussed here.
This, to me, is the mistake that SO made. Developers helping developers is the engine that runs the site, that's why people come.
The body of interesting knowledge is an emergent property of the support forum/peer-to-peer teacher.
Eventually they tried to put the cart before the horse and traffic has dropped.
It's OK that you're done answering "how do I change font color with jQuery" for the thousandth time and are only interested in the occasional very interesting question, because there are people behind you who do want to answer that question. That will help them grow to get where you are.
If we don't allow new generations of users to go through that process we went through, then StackOverflow has an expiration date.
ouch. as a former user of 10 years, I felt that. so true.
I think ChatGPT is amazing foe this because you will rarely have to deal with anyone who has this type of attitude when you need help in the future
The same mechanisms that make SO kind of brutal have also helped revolutionize asynchronous online Q&A.
I have asked a lot of questions, got good and useful answers on some. I have also answered a lot. Some of my answers and questions have been edited afterwards, most of the edits where good.
My experience was not perfect. Some of questions have been marked as duplicate when in fact they were not duplicates, but only superficially similar to other questions. Some of the edits people did to my answers where incorrect.
But overall, my experience with the site is still mostly positive. It saddens me to see that others are not enjoying the site as I did.
"Whats is wrong with the api?": moved to github, because its the source. "What api to use?": stackoverflow is still the place to go (or chatgpt). "How to improve api usage/syntax?" stackoverflow / chatpgt.
Github could really dominate with its own LLM in the web and Stackoverflow could regain some shares with an integrated LLM.
PS: I don't like Github Copilot because of UX and code upload.
But for everything else, it's mostly filled with folks who are not smart enough to comprehend the question and therefore often answer some other thing which they've pattern matched.
ChatGPT 4 can do that instantly instead.
Try asking a few programming questions and SO won't be the top link anymore.
__________________
0. https://www.zdnet.com/article/stack-overflow-ceo-on-how-it-b...
Stopped being active then. That was a few years ago.
The reputation thing is still nonsense. I can ask a noob question and end up with 10K reputation because of the likes it gets or I can answer a complicated question that takes all 25 years of experience to articulate well and get nothing. People can offer a bounty but it should also be possible to "reward" people who obviously know what they are doing and not reward people for asking questions.
The problem with duplicate questions is big, but again, the search isn't great at finding them, especially if you don't quite know how to ask the question.
They don't do much about people who just appear with 1 reputation asking something like "What does Null Reference Exception mean?". I ended up being allowed to moderate questions but the moderation queues get stuck constantly so I couldn't even help with that - after all, the crowd-sourcing is a great way to solve the volume of moderation needed and some of us are happy to help but as soon as you start seeing "Moderation queue full" after only moderating 2 posts, then you give up.
https://stackoverflow.com/users/1002260
for FIVE YEARS, because of a rollback war on some of my own answers. some idiot was gaming the system and adding minor punctuation here and there, just to get review points. so yeah, good riddance.
I have since deleted my account on SO, which was actually a somewhat involved process. They did not however allow me to remove the answers that I provided.
There was a chatGPT research paper where it could learn from its mistakes and seek novel solutions.
The first generation of SO people were folks who cut their teeth pre (or early) internet when it required quite a bit of effort to learn a technology. Mistakes without googlable solutions, patchy documentation, bad or non existent internet forced people to work hard to figure things out that it's in that effort where, in my experience, learning happens.
These people contributed to the early SO and genuinely enriched it. Made it a source of high quality topical information. Over the years though, it's become a source of cut/paste code and then perhaps cut/pasted code (driven by the gamification). Few people go there anymore to contribute well written deeply thought out answers. It's mostly fly bys I imagine.
I used to train freshers and over the years, I can distinctly see the decrease in quality and the increasing tendency to have SO, ChatGPT, whatever just "solve this problem for me" rather than "I want to learn this and get better at my job".
Sure, times change, but it's definitely not the fault of users that a platform is becoming obselete.
This like a chef saying: "People in the past used to like it" in response to anyone criticizing the meal
I think, SO, for the most part was a genuinely well intentioned effort with really good outcomes.
Spolsky's earlier talks describing his philosophies about a good QA site, I thought, were insightful. To their credit, they did organically reach the top of Google rankings and for a long time, the quality of answers and insights on the site were really good. I learned a lot of from reading answers to various questions by Alex Martelli and Raymond Hettinger (of Python fame).
I'm sure they made some mistakes along the way and those contributed to the current downfall. My larger point, however, is that this and other tools and sites which make things "easy" have contributed to a decline in the quality of (especially new) programmers and that has indirectly taken a toll on the site.
To extend your chef analogy, the rant would be "Here I am trying to create interesting dishes with modern ingredients but people these days haven't tasted really food and can't digest anything other than a quick burger, canned soda and frozen fries."
I can tell you that I quit in the last 6 months. All new answers are answers on questions from first time posters, answering questions that have good solid answers from 8+ years ago (mostly 10 years), and it is pretty obvious, they are just doing it to get points. The question/answer is not something that has changed in the last 8+ years, they are just doing it to get reputation or whatever.
Most good solid questions...I just do not see too many today. I am not saying this is the reason for the decline, but there is a deluge of posts answering very old questions with a slight modification of an old answer. I quit putting effort into it, as it was just non-productive.
It's entirely possible that the same is happening to SO.
One of the things I say with some frequency is that all failures are engineering problems. Blaming the users/customers, in the end, fails to take advantage of an opportunity to improve the product, learn from mistakes and improve process.
While I understand it's a lot of work, and some (most?) might be posting low quality answers, my main problem with stack overflow has been exactly that over time, the top answer becomes stale.
The question is still relevant, but the best practice or library, platform whatever changed... That is a problem the stack overflow model has a hard time adapting too.. and it just gets worst over time..
Original accepted upvoted to +20 answer: "Take Y and drive it into X, grep output for Z and place in AZ for processing with the -kombucha switch"
Answer posted yesterday by "NewUserXX123", a first post: "Take A and drive it into B, grep output for C and place it in DD for processing with the -kombucha switch"
It gets pretty tiring responding with a custom comment and downvoting, only to have "NewUserXX123" then cuss you out in a DM about how you are a total @!#$%()&!@$%.
I hit upon those useless answers all the time. I have no need to know about solutions that might have worked a decade ago and are obsolete. Stack Overflow is useless now.
Gamifying reputation was one of the worst things we've done. I was part of this in previous positions and after a while it became apparent that it brought out the worst in people. There are better ways of incentivizing usage rather than stupid internet points.
Agreed and I don't know this for sure but the points lust, it seems to me, is not limited to questioners and answers, some of the behaviour I've seen from mods seems to be driven by that too.
Just to repeat, I don't know that because I don't know what gets you points but the behaviour is hard to understand otherwise.
I had a question where I had finished with a one word sentence "Thanks." . A mod had removed that part of the question, when I looked at their record they had an awful lot of edits which consisted of similar, tiny, non-substantive changes. It did make me wonder about their motivation.
Anyway, yes, imho gamification not a good idea.
https://stackoverflow.com/help/behavior
"Do not use signature, taglines, greetings, thanks, or other chitchat.
"Every post you make is already “signed” with your standard user card, which links directly back to your user page. Your user page belongs to you, so fill it with information about your interests, links to stuff you’ve worked on, or whatever else you like!"
"Thanks and other statements of appreciation are unnecessary, and, like other chitchat, should not be included."
"If you use signatures, taglines, greetings, thanks, or other chitchat, they will be removed to reduce noise in the questions and answers."
----
Removing it doesn't get the editor any points, they are spending their time cleaning up your question to site standards for no reward at all.
https://meta.stackoverflow.com/questions/260776/should-i-rem...
> Wolfgang Amadeus Mozart[a][b] (27 January 1756 – 5 December 1791) was a prolific and influential composer of the Classical period. Despite his short life, his rapid pace of composition resulted in more than 800 works of virtually every genre of his time.
it opened with:
> Hello everyone, I studied classical music in college but have been out of the scene for a while and now I'm getting back in, I was wondering about the history of Mozart and I remember that his middle name was Allen or Almond or something? Can anyone help? Thanks for any information. || Hi, I learned about him in middle school and we all joked that his middle name was Armadillo but I think that's wrong, haha. || Greets all, it's in Olivier Hallengrunsch's Classical Composers as 'Amadeus', and that's well regarded. It also says he lived (27 January 1756 – 5 December 1791). I think we can all agree he was prolific, right? Thanks and regards, Jason [xxKiller; AMD Ryzen2 32GB RAM 2TB Western Digital SSD; BMW 330 2l aircooled] || Why's nobody mentioning how short that life was, smh || Good evening all and sundry, m'lady (tips hat), forsooth would anyone speak to how many works he is believed to have composed, all considered? Methinks such knowledge would be a most hearty addition to this esteemed gentleman's biography - Martin, [Fort Lauderdale TX Ren. Faire organiser 1997-1997] || etc.
StackOverflow isn't a forum, it's a collaborative reference work. Meta-chat would be edited out of a wikipedia page and goes on a separate 'talk' page (equivalent: meta stackexchanges or the stackexchange chat). What if you then went to the Wikipedia talk page and said "Is overzealous moderation of questions ruining Wikipedia? I want to be able to edit questions and greetings into pages but there are hoards of awful literal-minded jobsworths cruising the site just looking for the slightest reason to edit a spelling or grammar mistake or revert my changes. It just leaves me feeling unwelcome"?
Why would you expect to feel welcome when you're spoiling what others are trying to build up and insulting them for caring??
Some really old questions (circa 2009) have the answer to the latest version, that required scrolling through multiple pages.
I'm just not sure what value there was to have a team of people volunteer their time to making sure there were absolutely zero duplicated questions/answers? What problem is that solving? If anything, it is adding problems by not letting the content naturally evolve in time.
If people want to assess the quality of answers, it's better to have multiple data points, and SO could have invested those resources in algorithmically linking similar questions and making it easy to navigate by both answer reputation and time.
Along similarly lines, it seems best to tackle the people trying to game the system algorithmically as well - if content is word for word duplicate, that's a problem which can be solved by computers instead of people (similar text is a solved problem).
It seems the only use case for moderators on SO is for removing truly inappropriate content - it's wild to me that moderators were spending significant amounts of time actually removing technical questions and answers.
It reminds me a bit of reddit moderation, where reddit communities enforce these non-sensical rules and by extension require hugely heavy handed moderation to 'curate' their communities. Like, the headphones subreddit disallows pictures of headphones in boxes. Why? If the community is interested in headphones, what's the difference if it's a box or not? It's not like if you take it out of the box the headphones look different than any picture you can find on the internet.
Seems like many of the problems of moderation are artificial rules endlessly being enforced by real people which ends up just being pseudo 'make busy' work.
It would more sense to establish lineage. Lock older questions from more answers after some point in time, and instead of closing and linking to original, do it the other way. Keep new questions open and link to past variations. Then in the old threads indicate newer guidance may exist and link forward.
IIRC, the vision of the perfect question was one with a canonical answer. They didn’t want to be Quora. I think that concept made and makes a lot of sense, but doesn’t capture the “meta” issues surrounding how you litigate the form of the question, especially as questions get more nuanced.
SO has been a great site for me. Not too long ago, it was always my first stop, when looking for correct solutions to difficult problems.
I'm a highly experienced engineer, and a first-class debugger. I always get my bug ... eventually.
What SO would give me, was a correct solution, very quickly. I would ask a question, and have two or three excellent answers, in a matter of minutes.
I could definitely have found the answer, myself, but I would have had to do stuff like set up playgrounds, or even full applications.
Nowadays, they require question askers to do exactly that. In The Days of Yore, I could ask a question, without having to have gone through an hour of debugging and prototyping, and have a great answer.
But, at its heart, it's another wiki documentation site, and wikis don't age well. The WordPress Codex is damn near worthless, and even the great PHP docs are showing their age.
1. Copilot and LLM trained on code is reducing usage of StackOverflow.
2. People are just coding less and searching for less with layoffs and slow down due to burnout.
But it's definitely connected to these points. The degree of magnitude for either though is difficult to assess.
If it is mostly #1 though that's the worst thing isn't it? Because SO is weakened and then the LLMs gradually become less effective. We generate less curiosity and public discussion overtime as people use AI and LLMs more. That's sad for software engineering as a whole. One of the best parts of this job is collaboration and problem solving. Not just individually but as a community as well.
What I see recently is that the answers start being out of date, and Google sometimes bizarrely returns these weird copies of StackOverflow instead of the main site. Can it be that?
Also ChatGPT has really been a big help, as for most basic questions I can just ask it and get a fair result. This is more tricky to do when searching on google.
I.e.
chatgpt: How do i sort a slice of structs in golang with an inner field of "Name" Usually somewhat correct the first time.
google: golang sort slice struct (I hope to find a good answer, but I usually have to go through 1-3 pages
This is more documentation based, but this is also what stackoverflow did in the past quite well, that is answers to common questions for a language.
I guess I’ve always interacted with SO in a relatively passive way. There’s just too much toxicity that’s plausibly deniable. And with it, lost much of its usefulness.
Every time I try to look for an answer against knowing better, I still get disappointed.
Either the question isn’t there or someone asked it but it was marked as duplicate with a link to a decade old (and for some reason never more recent) question that is vaguely of a similar topic if you squint hard enough.
It’s silly anyways because some of the frameworks I work with are extremely volatile in the sense that they fundamentally change from year to year, so the chances that an older question is stale are all but guaranteed.
I don’t even bother asking the question myself because I see the copious amounts of pedantry while I browse.
Never mind discussions about “the right way” of doing things, I’m talking about entire comment threads about the right terminology of the “well akschually” variety.
Think arguments about whether something should be called a method or a function or whether MVVM is possible with framework X because the author had intended X, Y or Z, when none of it is even remotely relevant nor helpful to the question at hand.
Other forms of pedantry are entirely rewriting the OP’s question because a variable name or function name was disliked, even though the original question as asked was not confusing in the slightest.
The single time I asked a question, it was genuinely challenging and relatively low-level matter, or as close to it as you get in my field. Crickets, aside from 1 extremely low quality answer that didn’t actually answer anything that is.
I tried to focus on answering people’s questions instead, especially on a few relatively new frameworks that come with new conventions as those generate a lot of questions and I managed to become really proficient in them. I figured it would be a good way to give back to the developer community as a whole since it was that very community that enabled me to become proficient in it in the first place.
But stopped doing that too because it’s useless. Either good questions get closed, mods fuck around editing things around or the high level clique just upvotes each other’s low effort drivel.
SO is all but completely useless to me, it has become my last resort and only if I’m really stuck, and every time I dread it like I dread pulling teeth because every time it proves to be an exercise in wasting my time.
For me, there's no incentive to answer questions. It takes too long to get an answer, and the reply is usually some snide comment because too many assume it's an XY problem. It's generally worse than asking ChatGPT even when the AI gets it horribly wrong.
As for answering questions, I'd rather do that on my blog. Posting to SO is a time investment with zero return.
I was surprised to see big humps over early springtimes. As if springtime causes people to feel more willing to open up and communicate, similarly to how we associate springtime with heightened romantic drive.
Another cool thing I've discovered by reading this right after the Usenet thread[0] is that it is typical for large-scale social networks to stop being humanity's darlings after ~15 years of age. Which is also surprisingly similar to how people are supposed to become completely self-sufficient around 16, as if there's no reasonable expectaction of substantial external care for them.
Usenet: born in 1979 [u1], stagnated in 1993, 14 years later. XMPP: born in 1999 [x1], stagnated in 2013 [x2], 14 years later. Stack Overflow: born in 2008 [s1], stagnation reported today here, 15 years later.
[0] https://news.ycombinator.com/item?id=36859510 [u1] "Newsgroup experiments first occurred in 1979." https://en.wikipedia.org/wiki/Usenet [u2] "Segan said that some people pointed to the Eternal September in 1993 as the beginning of Usenet's decline" https://en.wikipedia.org/wiki/Usenet
[x1] "released the first version of the jabberd server on January 4, 1999." https://en.wikipedia.org/wiki/XMPP#History_and_development [x2] "In May 2013, Google announced XMPP compatibility would be dropped from Google Talk for server-to-server federation" https://en.wikipedia.org/wiki/XMPP#History_and_development
But I stopped contributing seriously maybe 2 years ago. I just got so sick of SO behaving like complete jerks again and again, and lost any desire to keep enriching their shareholders.
But yes, you're correct 100% might be a bit over the top.
GPT written code gets committed to github (After human review, so its working code only). Github data gets shared with OpenAI (Microsoft). So GPT continues to improve.
The recent deterioration is due to reduction in parameter count for GPT-4 to save money.
Yes, I agree that isn't what we're seeing now. But we will see it.
It's the hamburger feedback but without any buns around it. And I just wanted to contribute, on my own time often, to help somebody else with the same problem...
SEO is another factor. At least I don't get as many results on SO as I used to. That may be because of many reasons but if the popularity of a site starts to fall Google search might speed that up.
EDIT: I see now that the drop started long before ChatGPT became relevant. But it won't help SO I think. Github copilot and Google featured snippets may also be factors.
I especially enjoyed my time spent on various Discourse servers, e.g.
Figma: https://forum.figma.com/
CodeMirror: https://discuss.codemirror.net/
Compared to SO, those single product servers have much better signal to noise ratio and the community are very welcoming.
I firmly believe that ChatGTP is the SO killer. You get your answer instantly, it's usually good enough, and you don't have to worry about a mod closing your question as low quality or duplicate.
I did have problem like this before too and I just made my search better by appending the `after:<insert year here>` to my google search to ensure the latest info
Back in the day SO would be practically the ONLY place you would ask any questions regarding issues in programming or tech, but over time things like asking for help about specific frameworks or open source products has become more common to ask it on those projects github or discord.
And then you got LLMs being able to summarize an entire manual for you.
Yes, AI coding assistants will mean less frequent Googling occurs during development, but also, when Googling is required it usually goes beyond SO and more towards primary sources as the former's results strongly overlap with the AI assistants capabilities.
Stack Overflow has no good mechanism for ousting obsolete answers in favor of the current right answer.
The fun thing is... I think SO saw that coming. They made that huge push to be the central documentation hub for everything from huge projects to internal teams. It just wasn't well thought out and expensive iirc, and had an abysmal adoption rate as a result
My teams first go-to is ChatGPT as well. SO is no longer 'that site' to go to ( been on SO for 10+ years ). I have 0 love for that site. It was 'The resource' before, now it's just another site..
For more complex problems that are better served by a call-and-response loop, Discord seems to be where conversations are happening.
For me, I would say over 5 years ago I would post a question and get a flurry of activity within the first 10 minutes.
Now I post and I get no activity at all, and I end up answering my own question a day later.
Instead of being a platform that encourages those with experience to mentor and lead, and those without to seek experience, without punishment, it takes the opinion that “dumb questions are the ones that have already been asked and it’s the responsibility of the newbie to know if their question has already been asked.” It seeks to be the training set for its replacement rather than be its replacement.
Then a new product comes around to support the core use case newbies want more directly:
Yeah, of course we fucking left. Because surprise, if you’re constantly learning, you’ll always be a newbie at another thing. If SO had been an actual community, it would merged with an LLM rather than being eaten by them.
They usually seem to only know magical incantations that will get things to work and don't have deep familiarity because really competent people aren't googling for answers like we are so they never land on the page to give their 2¢
There needs to be different exclusion and barriers but not through the wacky constitutional sheriff system they are using.
Mailing lists are decent because you know you're going out on blast so social norms kick in. If you mail say LKML, most people imagine important kernel people taking their time to read your blabberings and that likely stops those who can't improve the silence.
SO OTOH seems too unbounded and freewheeling where popularity is more important than responsibility. It's the wrong cadence between invitation and expectation.
Anyways, grumble grumble.
This wasnt an issue I remember in the past. I've just kind of drifted away as a result.
Use SPA for the best UX for a forum. Forget the dumb Google SEO.
It was not sudden either, I remember by 2014 it was mostly pointles to try to ask any complex question.
Many simple questions can be answered in chat, and there are active Discords for many things.
Personally I hate it, but it is popular.
Unfortunately anything popular is a target, and once they stopped trying to innovate to keep ahead of systems gaming they were going to lose eventually. It's amazing it took this long.
but i suppose it will still be a very valuable dataset to anyone looking to train a coding model. its the single most valuable resource that could be archived for one
For more compelx problems, Stack Overflow was never a good solution; it's better suited for tackling simpler questions. But for simpler questions you have chatGPT.
I have watched/experienced SO go from a super useful/helpful site with a balance that pitted people's desire to be recognized with a desire to be helped, and slowly be subsumed by those overbent on organization and bureaucracy. And is it did so, it became less useful, to me as a questioner as well as an answerer.
Now I have to put up with stochastic parroting from GPTs to try and steer me in the right direction. Yay.