Google rolls out algorithm change in the US
googleblog.blogspot.com
googleblog.blogspot.com
It would be interesting to know how Google determines, in an automated perspective, which site is the original and which is the copy - especially when copying could go both ways. For example, I could write an article, license it under the GFDL, and someone could copy it to Wikipedia. I might then copy the Wikipedia improvements back to my site. Technically, I had the content first - so would Wikipedia be penalised?
If there is a bias against smaller sites, this might make smaller sites be reluctant to license their content under licenses that let bigger sites copy them.
But Wikia will refuse to remove the original wiki even if all the contributors want to move it, in order to maximize advertising revenue. This means that it's nearly impossible to get the new wiki to rank highly on Google, even if all links across the internet are changed to point to the new wiki, because Wikia's pagerank is so high that Google deems the new one to be a "copy". This is why Wikia moved all their wikis to subdomains: in order to piggyback on the pagerank of the main site.
As a result there are a ton of long-dead wikis on Wikia that still get more search traffic than the active equivalent. Obviously this hurts users, since they get long-outdated information as a result.
In short, once a wiki is placed on Wikia, it's basically impossible to ever move it anywhere else because of Google's anti-duplicate biasing.
See http://en.wikipedia.org/wiki/Wikia#Controversy for more info.
The sites Google seem to be targeting are those who aggregate content wholesale from a number of sources. One way they could identify these would be to examine the number of different sites from which a particular site appears to have copied its content from.
In any case this is a bold step for Google. It shows they still care for the quality of the search and have guts to take big bold decisions to protect it.
If a site has N copies associated with it, then you could compare that to the average number of copies associated with sites of that pagerank. If the difference between N and the average is high, then that's a likely spam site. Let's call this difference σ.
This gives us a problem though, because while σ is a good indicator of spamminess, it's not foolproof. A low pagerank site could have been copied from lots of times, which would unfairly earn it a high σ.
What we can do is calculate a σ', by inversely weighing each copy with the σ of the site on the other end. A copy of a site with a low sigma will increase σ', and a copy of a site with a high sigma will have less of an effect on σ'.
So while our high σ may have initially suggested the site to be spammy, all its copies are from high σ sites, which are counted less towards σ', leading to an overall verdict of not spammy. [I wonder what would happen if you iterated this process]
As for your example, the smaller site wouldn't be penalized, because N would be low. If it were say, a full wikipedia mirror, then it would be penalized. Wikipedia would not be penalized, since it has an ungodly amount of pagerank. There's also no bias against smaller sites, since σ is calculated relative to pagerank.
What have I missed?
Does anybody else feel that Google has been figured out now, and that they are now just applying patches to a system that is fundamentally broken (in the same way Excite and AV did)?
There might be an opportunity here for a new search engine - one that creates a new core method of ranking, as PageRank was, to filter out duplicate content, content farms, malware sites, etc.
"Relevance = links" today looks as tired as "relevance = keyword density" did 10 years ago, but I have no idea what the new "relevance = " is
I don't. I think they do a much better job of actively improving their offering. No matter what the system is, someone will always try to game it. Even if the people running the system are very smart and try very hard, the smart gamers will (briefly) succeed. That's ok, though - it pushes the people running the system to work harder to improve it. It really is in Google's best interest (IMO) to return great search results, so I think they'll continue to improve what they do.
This Fallacy was first described to me in Brenda Laurel's Computers as Theatre. http://c2.com/cgi/wiki?BrendaLaurel
Also, I am not sure that the spam problem is as bad as people claim it to be. I have looked at virtually all of the lists of content farms that have been published since Google release their Personal Blocklist for Google Search and it seems to be that (a) it's all very subjective and people rarely agree that a particular site is a content farm (b) and therefore the number of sites that people seem to agree are farms aren't more than a dozen.
The social graph is essentially a personal thing, not professional. Assuming that we see a need to partition our professional and personal lives (and that should be true, in order to protect both sides), then the professional side isn't going to be reflected well in the personal graph.
So long as our networking tools don't recognize this split, they are going to be (relatively) starved for content that is strongly oriented towards the professional world. For example, I wouldn't expect my wife to be able to find deep information about Medicare reimbursement if we relied on mining social sites.
For example, I work at a healthcare services start-up, and we have one person full-time researching our competitors and figuring out how to differentiate us in the marketplace, as well as figure out what keywords someone might search for to find this kind of product. We're mainly mathematicians and smart programmers who found a way to break into the market initially, not marketers, so our terminology doesn't match domain experts well.
Anecdotally, I think most users are VERY bad at ranking how valuable information is. For example, we'll get paid $100,000/year for something that takes 1 week and provides zero insight into their business, and then they will refuse to pay us at all for something that makes them $30,000,000 a year! This is so common that I am wondering if its more a symptom of human nature than our customers.
Often times hospitals receive $0 for something they should be getting $10 million per year for, just because nobody ever contests the NOPAY response. Finding that sort of error isn't easy.
> It’s worth noting that this update does not rely on the feedback we’ve received from the Personal Blocklist Chrome extension
>However, we did compare the Blocklist data we gathered with the sites identified by our algorithm, and we were very pleased that the preferences our users expressed by using the extension are well represented. If you take the top several dozen or so most-blocked domains from the Chrome extension, then this algorithmic change addresses 84% of them, which is strong independent confirmation of the user benefits.
There's a difference between training with that data and testing against it, which is precisely what I wanted to highlight here since many expected Google to use such data for training. You can't do that since you can overbias/over-fit given what little daat we, plug-in users, are able to produce versus the amounts they have access to.
2nd edit: it should also be noted that that 84% figure they mentioned related to the extension data is a measure of recall. Assuming Google wants to do a better job with precision in these cases (don't want to get many false positives), that recall is still fairly good in mind since I'm sure there's some noise in that data.
There seems to be enough "normal" spam in the search results that Google should be focusing on first. Just yesterday I searched for "viking dishwasher clog" and the #6 result is a .info page that is nothing but obvious spam.
I previously mentioned the issue of Wikipedia poorly regurgitating content from original creators (sometimes with attributing, sometimes without) and outranking that original page, even if it contained approachable, illustrated and in-depth article, and Wikipedia's content was much weaker. This is not something that tech bloggers will notice - they are much more likely to read Wikipedia's programming, science, math, or tech articles, which are of much higher quality than those on other topics, so you won't hear them complaining when they see Wikipedia as #1 everywhere.
Another bias I starting seeing lately is the high ranking of new Q&A sites (another focus of tech bloggers), when excellent topical forums with very good content, often exactly with the answers I needed, are ranking poorly. I’m speculating here, but I think Google’s focus on links might hurt its ability to bring up deeply hidden content from forums. These forums won’t get many links from tech bloggers and others looking for the next big thing, but they have very valuable content for many long-tail searches.
I remember an example provided here was "nstoolbar bottom bar". Let's take a look: http://www.google.com/search?q=nstoolbar+bottom+bar
Granted, it was probably linked from HN somwhere which will have made matters worse.
If you take the top several dozen or so most-blocked domains from the Chrome
extension, then this algorithmic change addresses 84% of them, which is
strong independent confirmation of the user benefits.
Well, what if they are mistakenly penalizing Common Jack's blog because it gets re-posted across the interwebs? I bet no one is manually blocking Common Jack's blog, so that factor is lost.Example: let's say Google just randomly penalizes 84% of ALL domains. Chances are that they intersect with, you guessed it, 84% of the Chrome extension's data. Independent confirmation!
Small example: When you search "VLC" on google on find scams websites shipping VLC with a toolbar/crapware/malware, buying or not adwords...
Google plainly refuses to remove those websites, because they gives Google some money. So they make money from scammers, know it clearly and do not want to do anything...
And as we are a very small team of volunteers, it is impossible to attack those websites...
NB: this happens with other software too.
You may be right about what's happening, but I don't think it's fair for you to assert as fact what you believe their motivations to be. Matt Cutts and others have categorically denied these claims, and provided good reasons.
I very much prefer DDG these days, but those statements are just not true. And I do not believe that Google would purposefully provide _much_ _worse_ results for US customers, relative to EU ones.
Try .fr, .it, .es and you'll see...
And on .us, just deactivate your adblock.
The idea that most people really care about the purity of their search results, let alone would understand what all this means even if the BBC or NYT ran a proper article is just unrealistic.
Personally not having heard of it before, I think this comment (accidentally?) contributes somewhat to the discussion. Has HN really become this trigger-happy on the downvotes?
They have little original content, but are still a useful source.
As far as I understood the whole Bing ordeal was that users with the Bing toolbar reported not the links and ranking shown, but what the users chose as the "correct" hit for that search.
In that regard this doesn't really have to alter Bing's results in any way at all.