Stack Overflow Creative Commons Data Dump
blog.stackoverflow.com
blog.stackoverflow.com
It only took a few minutes to figure this out but if you're as confused by this quirky archive format as I was, there you go :) Don't bother trying to unpack the ZIP file in the normal OS X way as it'll just keep unpacking over and over and not give you anything useful.
It's one of the first things I install on any OS X system. Great unarchiver that takes care of most (all?) the common archive formats, although I run into the occasional password issue. It beats StuffIt without a doubt.
If spamxyz.org violated the cc terms would that be enough to complain to Google?
In a recent court case in the Netherlands, some company A filed a complaint against a website because it ran a story about another company B that went bankrupt and mentioned the plaintiff in an unrelated story on the same page. Searching for the plaintiff's company name and "bankruptcy", Google would show you a summary that looked like company A had gone under. Does that make the website responsible? Is it Google's fault? The judge decided the website should take responsibility and fix it. I think that you're wasting your time when you're searching Google to find out whether your company has gone bankrupt.
In the case of spamxyz.org: report it as spam. In the case of the Dutch court case: use your common sense. In the case of newspapers crying about summaries in the search results: use robots.txt. But people should stop pointing at Google to fix all the problems on the internet. They could do a lot better, but there are plenty of scenarios in which you don't want Google to do decide on their own whether they should show a site in their search results or not.
PS: regarding Wikipedia data: http://en.wikipedia.org/wiki/Wikipedia_database
http://download.wikimedia.org/enwiki/ http://en.wikipedia.org/wiki/Wikipedia_database