In the mean time, use this extension to clean up your own search results and tell us which sites you don't want to see in Google.
In the mean time, use this extension to clean up your own search results and tell us which sites you don't want to see in Google.
(AKA great stuff!)
Perfect example of this extension in action: search for "how to remove ear pads from HD555", and I get a link to this page as the third result: http://www.fixya.com/support/p516634-sennheiser_hd_555_consu...
Totally useless!
While yahoo answers etc. might provide bad answers, experts exchange provides no answers at all. All you see is an open question with a nasty subscribe button that will only work if you haven't already used it more than 30 days ago.
So experts exchange is more than useless: It's just plain advertisement without any added value. It's the kind of stuff you expect in the ads section of a search engine, but definitely don't want to show up in your search results.
I don't mind their business model, but I can't understand why they appear near the top of so many search results.
Either they are very good at (mis-)using SEO techniques, or there are really many websites linking to them. However, I personally find it hard to believe that any author of a blog article or forum entry links to an EE answer voluntarily.
ExpertSexChange
If we get a good signal from this extension, or from offering block links in Google's search results, then it's much more similar Gmail's spam algorithm, where an email is labelled as spam partially because a lot of users say it is, rather than because of some editorial decision on our part.
For example, sometimes we see copies rank higher than originals. Why does that happen? Google know where they first saw a particular piece of content, don’t they? Why don’t they use that as a heavy ranking factor?
Or am I too far off?
Now which is the original from Google's point of view? Relatively smaller blogs can take significantly longer to index then sites that have massive amounts of content moving about daily.
My small negligible personal website notifies Google, Bing and Yahoo immediately and automatically as soon as I publish something new. It also publishes a feed. Even if the content is picked up and republished right away by a site that is indexed every minute, it should be possible to determine correctly the original publisher.
In some cases, I can think of more ways to determine the original publisher. And certainly Google can think of even more.
Please share them. Thanks.
I publish an article at only-original-content.com. The article has some images that are served from only-original-content.com/images. Now only-copied-content.com takes my original article and republishes it. Since only-copied-content simply copied the HTML, the images are still served from only-original-content.com/images.
In that case it should be simple to determine who is the original publisher. Of course, only-original-content.com could simply be a CDN that only-copied-content.com uses for its static resources, but, again, it should be easy to determine whether that is the case.
That said, if you wanted to share some examples where you're seeing copies rank higher than originals I'm happy to pass that on to the right folks. In fact, some of the right folks are already on this thread. :)
Personally, I'm glad that you're putting this out as a user-controlled thing. I like the fact I'm able to get rid of results that aren't necessarily spam or SEO'ed garbage, but I where I still know that I never want to see results from that site again.
However, if it is built in to the general results, could you also add a metric to Webmaster Tools showing the frequency with which your domain is reported. It would also be good if the blocking could have timeout period so that sites can be given a chance to improve their behavior, rather than just deleting that domain forever.
Getting users to do the work will help with the most egregious abuses. One thing that content farms are good at is coming up with answers for questions that don't have an answer available on-line. For example, the query "what's the personal cell phone number of <insert celebrity>?" will just take you to a spam farm because that's the only kind of website that claims to know the answer.
We've absolutely done similar things in the past, but that comment was the spark behind this most recent Chrome extension. I hope that's specific enough. :)
Unless you want to categorize everything Matt says publicly (whether here, on his blog or on twitter) as PR...
Will you be offering a Firefox extension?
After using this for less that 12 hours and I think it's fantastic. I also think it could be vastly improved upon. What if you added reddit style "up" and "down" votes to my results?
Maybe it sounds silly at first pass, but if all my votes get passed back to Google, your algorithms could learn from an incredible crowd-source treasure trove of knowledge.
A black list wouldn't just help you find the bad ones - but up-boats could help you suss the good ones.
There's the concern of DIGG style voting blocks, but perhaps even that could be detected by sufficiently sophistimatacted algorithms. (detection of pairity in voting that fell outside of statistical norms could be downgraded in quality).
I mean, sure, I'm already gaming the system in my head ... automated creation of accounts, automated up-boating (which I suspect happens on Reddit as well, despite being bad "reddiquette")
OK. Actually, maybe it's a terrible idea. Like DIGG with only down-votes. Seems like a potential treasure trove though.
Right?
Is your block list something that will be synced so that it is available everywhere you use Chrome and eventually just synced with your Google account?