I'm not saying that a clone will never be listed above SO, but it definitely happens less often compared to a several weeks ago.
To be clear: the webspam team does reserve the right to take manual action to correct spam problems, and we do. That not only helps Google be responsive, it also improves our algorithms because we get use that data to train better algorithms. With Stack Overflow, I especially wanted to see Google tackle this instance with algorithms first.
Because we don't understand what's hard, we think you're not really trying, and then we make up evil reasons to explain that.
I believe if people understood better the difficulties of spam fighting they would be more understanding.
Not necessarily. The rate at which Google refreshes its crawl of a site, and how deep it crawl, depend on how often a site updates and its PageRank numbers. If a scraper site updates more often and has higher PR than the sites it's scraping, Google will be more likely to find the content there than at its source. Identifying the scraper copy as canonical because it was encountered first would be wrong.
Sites that are the victims of content cloning have to be very visible and valuable, so maybe a little manual curating could be relevant.
> the Stack Overflow cloners could just make other websites
Not really? The point is not to tag the clones but to tag the original; everything that is not the original and that has copied content is a clone -- its name, domain or country notwithstanding.
the primary input to search engines comes from web crawlers...the idea of "first" when it comes to duplicated content is already difficult to determine, and (I would guess) it would get much much worse in the inevitable arms race if something like this were implemented.
But I'd like to say one other thing. Why is Google only doing something about web spam now after people have pointed out how bad things have been getting? Has anybody considered creating a small team to just oversee public perceptions of the search results and try to keep on-top of things in the future?
This happens for more than StackOverflow clones. Mailing lists, Linux man-pages, FAQs, published Linux articles, etc. all have clone pages that are obvious link farms (sometimes they even include ads that attempt to harm my computer) that rank higher than the "official" (or at least less-noisey) pages.
Ideally, I'd like to completely remove domains from result as has been discussed elsewhere on HN. Hopefully this upcoming push for social networking that Google has will reintroduce a better-implemented "SearchWiki" feature...
Comment #1:
> I've been tracking how often this happens over the last month.
> <snip>
> it definitely happens less often compared to a several weeks
> ago.
Comment #2: > I am seeing many, many more clone sites in my search
> results in the last few months
You can't argue against "things have gotten better in the last week" with "things have been getting worse for the last 6 months."For example: https://encrypted.google.com/search?hl=en&q=aws+s3+emr+p...
Result #4 at the moment is "AWS Developer Forums: Interactions between S3, EMR and HDFS ..." on http://www.hackzq8search.appspot.com/developer...com/...
What's sublime about this example is that:
1. hackzq8search is clone of AWS's websites amazonwebservices.com, aws.typepad.com, etc
2. hackzq8search is hosted on appspot.com, Google's App Engine domain
3. hackzq8search is over quota, so the site doesn't show any content anyway.
Yet this site was the top search result, beating out the site it was cloning, time and time again on my AWS/EMR-related searches this week.
The one mitigating aspect as that hackzq8search's URL naming scheme is easily decodable -- the hackzq8search URL includes the full URL of the cloned URL, so I can write a Greasemonkey script to extract the proper original URL.
I found a glimmer of optimism in that the site has been slowly fading in SEO-success this week: I complained about https://encrypted.google.com/search?q=aws+s3+security+sox+pc... on Thursday, but on Friday the hackzq8search Search Result was gone from the first search result page.
It's still not hard to slam some AWS-related keywords into Google and get these bogus results, though.
The efreedom answer at the 5th position is actually the most relevant - the stackoverflow question from which it was copied doesn't even show up on the first page. There is one stackoverflow result on the first page, but it deals with a more complex related issue, not the simple question I was looking for.
It almost feels like a cache miss when I have to drop down to the official site/documentation, since that typically requires a greater time investment to read through to find the relevant sections.
I guess that's a tribute to how well stackoverflow works, most the time. And also to how lazy I am.
Of course, then all the content-copy farms will respond by copying valid content plus word lists - hopefully Google knows how to detect that.
And then another SEO cycle would start. Don't forget that before google came along nobody was trying to 'game the system' with backlinks and other trickery, the fact that that google is successful is what caused people to start gaming google.
Typically you pretend the search engine is a black box, you observe what goes in to it (web pages, links between them and queries) and you try to infer its internal operations based on what comes out (results ranked in order of what the engine considers to be important).
Careful analysis will then reveal to a greater or lesser extent which elements matter most to it and then the gaming will commence. Only by drastically changing the algorithm faster than the gamers can reverse-engineer the inner workings would a search engine be able to keep ahead but there are only so many ways in which you can realistically speaking build a search engine with present technology.
Your ideal, I'm afraid, is not going to be built any time soon, if you have any ideas on how to go about this then I'm all ears.
Stackoverflow comes in at number 8 while clones are 6 and 10
The reason Q&A sites are so visible is that people tend to type questions in their search engines, so Q&A sites are a good match to those.
Moreover, the unique licensing around SO content, along with its mass, presents an interesting edge case for Google. They should of course fix it, but it's not indicative of the average or mode experience.