Folks, careful what extensions you add.
Personally, I'd look for a link from e.g. eff.org, or a similar trusted site.
Why should you trust me? I'm Larry Page. Why should you trust that I'm Larry Page? Simple. Look at the "Who am I" part of this post.
Who am I: Larry Page, obviously.
As a footnote, I'd never run something like this on my computer, but if it's just IPs you want, an RPi image would get you my IP.
And heck, any and all phishing sites, by definition.
So millions of websites?
That's weird, because they seem to have quite a lot of coverage in major tech news sites:
Nothing wrong with that, and asking the question is fair enough, but doubling down on it in the face of being given references to who they are that are easy to verify gets a bit tiresome.
Especially since this is a group prominent enough to have a brief but well cited Wikipedia page [1] with references to media appearances in Wired, The Register, PC World, Technology Review, OnTheMedia, Techdirt, BBC, RadioNZ, CBC, New Scientist and others.
You can also find Jason Scott on HN [2], and his comments will point you to a number of other HN threads discussing archival, some of which are relevant to Archive Team.
They point to the Archive Team in point 4
Check http://chat.efnet.org:9090/?channels=%23noanswers for updates.
Without the %23 prefix.
Sometimes there's some good content, but you gotta wade through the 99% sewage to see any of it.
Asking better questions is normally a "better" way to find higher quality content.
But there's less of that than one might think. It rapidly reaches the state where the questions become highly specific (e.g. diagnose a medical condition, fix a bug, solve a personal problem, homework questions), repetitive, or argumentative. The Q&A format just isn't very good for long-term growth.
So there probably is good content on Y!A. (I like to think I contributed some myself.) But that quickly becomes swamped under a vast morass of unanswerable questions and poor-quality answers.
Nobody seems to have found a good solution to that. StackOverflow seems to be doing best, but at a cost of keeping its community small (small enough to fit in a single rack) and highly focused. Having a continuously changing technology helps -- but on a more general Q&A site that rapidly degrades into arguing about politics.
The projection given on IRC was "There's no way we'll get them all" and "assuming 150 million items, we need to go ~5 times what we are now" so help is appreciated!
Aren't you also providing 100% of the resources if they are using your IP, or is there a way for them to do most of the work "behind" your IP?
100% of what resources? Bandwidth, yes - you need to download the answers and then upload them to the archive. Fortunately it compresses fantastically so the uploads are fairly inconsequential. Even more fortunately, Yahoo! Answers is text-based, making the bandwidth usage pretty minor compared to almost any other past Archive Team project. It's really the overhead of each individual network request that they're trying to outsource here.
But for storage? You're providing approximately 0% of that. Files are only stored on your device for long enough for your computer to hand them off to a more permanent home at archive.org. CPU usage is also very minimal - again, thanks to text compressing super easily, the compression settings don't need to be particularly aggressive.
Does anyone happen to know what safeguards exist to protect against random child porn or other illegal content from flowing through my connection?
I checked the site but didn鈥檛 see anything address this concern.
My fear is the archive grabs a page that happens to have nefarious content which I鈥檇 be legally on the hook for.
Am I being overly cautious or is there a genuine risk?
You're being overly cautious. There is no legal risk in viewing a Yahoo! Answers page (or any other ArchiveTeam Warrior recommended project for that matter).
I was worried about random websites that get archived on archive.org.
Alternatively, the collection code could do something weird or suspicious, but it's open source and the team has a good reputation/enough social proof leading me to consider that unlikely