Google Web spam - Gabriel Weinberg's Blog
gabrielweinberg.com
gabrielweinberg.com
Unfortunately I don't have a better way of assessing it to offer. Internally, we often look at impression-weighted precision as a metric, but I don't think there's an easy way we could expose that to you.
A more reasonable thing to do would be to take a sample of DDG's query logs, scrape the results from Google, then see what percentage of Google's results come from your spam domains, but that requires sending a lot more queries to get any useful data.
That said, if anyone has a meaningful sample query set, I'm certainly interested in running it against my spam index. I see a lot of hits on it via other search APIs.
By the same token, if you are trying to launch a new spam classifier that has some false positives, if one of the false positives is yahoo or facebook, it doesn't really matter how good it is, it will never be worth the collateral damage.
As a result, rather than measuring precision/recall as a percentage of domains or as a percentage of urls we usually try to measure it as a percentage of results that appeared on a search result page or results that users click on, mined from our logs.
This is one of the bajillion reasons why it's absurd to expect Google to throw away all its logs data. The logs are essential to coming to any meaningful conclusions about the current quality of our search, let alone finding ways to improve it.
Am I right in understanding that if a SERP result for a given keyword doesn't get clicked by users enough, it will be removed?
By the same token: does the result's bounce rate matter? I imagine spammy sites have a very high bounce rate.
So this basically means that you are able to discern content spam on authoritative domains (facebook, wordpress.com, etc) based on ctr, bounce, impressions compared to surrounding serp results rather than comparing that data against the parent domain as a whole?
It seems the Google search APIs have actually gone backwards over the last few years.
Google doesn't want the search market to be fragmented; they want to dominate the market. I think that's why Google doesn't have a good search API offering.
There's not a search engine out there that wants to allow you to muck with their relevance algorithm by changing the results, and Google has more to lose from DDG-like activity than it might gain.
That said, I wanted to acknowledge that this isn't ranking data. However, perhaps as a result of this post, I'll be able to get some and re-post those results.