Awesome point about the political blog example. Perhaps attaching some sort of meaning to the filtering instead of a black/white SPAM vs. NOT-SPAM would be an added advantage of leveraging human judgement. So for example spam on a political oriented site would include news about celebs; or spam on gawker.com would equate to political links.
So the levels of filtration could be:
1) SPAM
2) Not relevant (to celeb/political/geek/etc nature of site)
3) <suggestions please.../>
The problem with this is the notion of "purity" between a collection of social news sites. Too much homegeneity (is that a word?) is problematic as diversity is what brings value to social sites. I suppose having a level of how stringent the filtration should be an option for subscribing to such a database. There should also be a way to decentralize this as centralization is not good (just my opinion...).
As far as spammers setting up separate websites for each social site....GOOD POINT!! A Bayesian filter that would introspect different links in a database could help match different sites at the content level with some degree of confidence, thus exposing the level of sophistication a specific spammer and identifying higher spam threats. Similar content could be matched between different databases (if decentralized) and such information could propagate faster between different types of social sites identifying spammers by content.
Because human judgment would trickle down to a central location (how can this be decentralized I wonder...) means that the timing of the appearance of sites in the database would possibly be significant (wouldn't it?), adding yet another insight into how spammers might operate (ok, perhaps I'm stretching it here).
I think the best way to look at it is to have a human filter (that depends on mass collaboration) up front and then later use a Bayesian filter to associate at the content level the different websites. Creating a web crawler to do this would be pretty trivial (I would imagine).
------------------
How would spammers game such a system? One thing they can do is NOT report their own links as spam. This does not prevent others from doing so. Spammers could conversely report other sites as spam thus muddying the waters through volume. I can imagine a scenario where they would muddy the waters through their own sites, which would translate into merely blocking that site as a preventive measure (this is DEFINITELY fraught with problems that I have not thought of yet!!!). Similarly if they attempt to muddy the waters through a legit site, then there is the advantage of genuine human involvement and randomness.
This becomes an issue of membership to the database: what are the criteria to submit to the database and be a contributor to flagging spam sites? I suppose subscribing to such a database would be free for all.
A possible solution to centralized authority would be for everybody to contribute, but at the same time a site can allow to trust the quality of reporting from certain sites and exclude others; this could have yet another "emergent" quality of identifying collaborations of sites that wall off spammers trying to muddy the water: spammers can muddy the water as much as they want, but they won't be heard. This would eliminate the need described above for judging different sites; also, it could provide new sites for a basis of which sites to trust for authority. Theoretically the head of a long tail of trust should emerge, and hopefully this could possibly be a shifting mass of authority. This decentralization of authority would allow spammers to do their worst and still not have an effect (wouldn't it?).
But somehow I feel intuitively there might be a way to game the system...any thoughts???