I thought that this was a massive nono from Googles side, has something changed?
I thought that this was a massive nono from Googles side, has something changed?
But the crazy part is that, for example - Ahrefs says that StackOverflow has "Organic traffic" in the range of 22 million per month. A lot of these copycat sites, at least the ones I saw - have a traffic range anywhere from 10k to 500k per month.
I mean, it's pretty insane just how well such sites can rank in Google, and you bet those copycats are making absolute bank from ads even if the majority of developers immediately close the site.
There's a lot going on with Google Search these days, a lot of people are complaining that sites that scrape content can easily rank really well for long-tail keywords. One case in particular, a site will scrape Google to collect "featured snippets" and "people also ask" - then combined anywhere from 20 to 40 of these answers and publish them as a blog post.
None of the words are changed, all questions/answers worded exactly the same. And Google puts these sites on page 1.
What a joke.
yeah i've been hitting a ton of those lately.
Would they just move to creating and using new domains with the same content as soon as traffic to the old becomes drops? (What looks like the spammers in the original post are doing)
But something does need to be done to these sites.
This is a decades-old spammer trick. Google used to not rank brand new domains very high for this reason.
It's hard not to think that the only reason Google abandoned most of its old site ranking heuristics was that they were filtering out too many sites with lots of Google ads. The spam sites now infesting Google's first-page results don't look very different from the spam sites I saw back in the early 2000's. (There's more JavaScript, but modern search spiders run every page in a VM before reading the DOM, so that doesn't fool anyone.)
I bet the majority of developers block ads
It is simple. Google is making more money from copycat sites then from original content...
Actually forcing a search engine back into the reliable index of valuable sources would be great.
Imagine you to a white list approach to a search engine where a human or AI does an approval first.
I don’t think this is a cultural issue, I fail to see how this can be considered value add by anyone.
It blocks copycats and hide them from multiple search engines. You may also use the list with uBlacklist.
* the identical text copied from stack overflow should be easily identifiable
* volunteers put together a list of these sites themselves
it should be obvious to Google apoligists that Google is either negligent or intentionally allowing these sites in their search. I'm sick of hearing about how "the world is different" and it's an "arms race" between spam sites and google. Bullshit.
Google starts matching content from SO => Spammers start tweaking the text slightly => google implements some expensive similarity score to down rank copy cat sites => spammers use more complex scrambling=> ...
> volunteers put together a list of these sites themselves
These lists only work because they're used by a tiny minority of people. If Google were to do this the spammers would start switching domains more quickly (or find some other workaround).
I'm no Google apologist but I think you're underestimating how hard search ranking is when spammers are actively trying to game the system.
That's what ML is perfect at detecting, which is Google's forte.
Some of these sites have been returned as top results for a while, so are you suggesting that Google just gave up because spammers would be able to evade them with an update?
You underestimate the resources google has at its disposal.
They simply don’t care because there is no real competition to worry,even with this spam you are still likely to use google, so why would profit motivated company bother ?
At the very least they're being deliberately neglectful because they don't feel the bad experience harms their revenue because there's no other substantial competitor so they can abuse their monopoly status.
I guess they may just not care enough about software developers and figure we're mostly using ad blockers so its wasted effort and we'll develop blocklists ourselves. With no monetary value that they can assign to the ill will that it engenders they figure it must not matter so they don't bother. Pissing off a large chunk of the entire IT community via obvious neglect seems like a poor move to me, but then I've never felt that I'm cut out for management.
It feels like economy-wide that decision makers in corporations and governments have just arrived at the conclusion that there's no money / no point in trying to stop scammers (and there might be an actual cost to revenue of doing so). It won't goose their quarterly numbers and might hurt them so its better to allow it.
This changed with the Google "machine learning" days, where you no longer have humans at the helm laying down explicit rules, so no more "change the world" updates, you can only slightly nudge the parameters towards what you want, meaning the same old tricks keep being effective for far too long.
That's just what the scheduled "core update" days are now: https://developers.google.com/search/blog/2022/05/may-2022-c...
A lot of updates are targeted at specific problems such as low quality product reviews but there are still broader updates taking place.
Maybe that means we should be searching in yahoo rather than google.
Just go to SO and use its search bar. It's actually quite good.
I mean, you know that's where you'll want to find the answer anyway - not some random corporate webpage or ad-infested splog. Why not cut out the middle man?
Only if that fails do I bother with Google.
I think a lot of others formed their opinion (myself muchly included) about this from sites where the search bar was a joke played on people.
Edit: let me upgrade that 'fine' to 'great', now that I think about it it was actually better than a google search which was not my previous experience.
I've found some fantastic articles out there, yes SO is a fantastic resource but there is an entire internet out there :-)
10 years I gave up on a large project where I rehosted and organized dead Usenet forum content because Google's dupe-penalty detector was too good and too aggressive for content that you could barely find beyond a six-year-old cache hit where the origin website was long gone.
Meanwhile these Stack Overflow scrapers are just `<html>{copy-and-paste}</html>` and the same domains are still alive despite years of cloning.
Looks like it's time to boot my project back up.
https://www.searchenginejournal.com/ranking-factors/google-a...