ArchiveTeam needs OPMLs and feed URLs to grab cached data from Google Reader
allyourfeed.ludios.org
allyourfeed.ludios.org
If you're interested in being able to read old posts in some future feed reading software, or just like having the data preserved, you can upload your OPMLs and ArchiveTeam will make its best effort to grab the feeds.
More details: http://archiveteam.org/index.php?title=Google_Reader
7TB+ of compressed feed text: http://archive.org/details/archiveteam_greader
Also, if anyone has billions of URLs that I can query, I could use them to infer feed URLs and save an incredible amount of stuff. See email in profile if you do.
https://github.com/epaulson/stash-greader-posts
It doesn't upload them anywhere, but at least I've got my own copy of them if I ever think of something I want to do with them.
Also, it looks like there is another tool mentioned in https://news.ycombinator.com/item?id=5958188
I have a couple of question though:
Will the data remain archived on my system after it is updated? And what format will that be in?
Will there be a public API to access this data once uploaded, or for services such as Feedly to import back entries from feeds? (I would hope they would support that, but the public API would be enough for me.)
Thank you for providing this service.
No, not currently. But the Internet Archive will provide the raw data and anyone is free to setup such an API :-)
Thanks for helping out!
As for an API, someone will hopefully write one to directly seek into a megawarc in that archive.org collection, or import everything into their feed reading service.
Among other things I would like to set up an ElasticSearch cluster for my own feeds.
Is the WARC format defined somewhere? I haven't looked at any other ArchiveTeam projects so I'm not informed if this format is used elsewhere.
There's an ISO spec for WARC and tools linked at http://www.archiveteam.org/index.php?title=The_WARC_Ecosyste...
https://github.com/ArchiveTeam/seesaw-kit/blob/master/run-pi...
Also, I think the backend is the hard part.
Your feed collection is like your personal life. It should be private (so what if the URLs are public and general). By disclosing your collection, it is like you are living in a glass house.
Unless I'm missing something, I don't see a single helpful reason for this service.
But why would ArchiveTeam wants to preserve the historical items in a feed if the feed does not belong to them in the first place (neither did it belong to Google)?
It's fine if people don't see any privacy implication here by submitting their reading collection. But as far as my single individual point is concerned, I don't see why I should upload my OPML for the sake of preservation. I have hard time uploading it to any other Google replacement out there trying to compete.
That's exactly my point.
There's also no telling if InoReader is open for grabbing what they've grabbed already again. Meaning it's potentially behind closed doors.
ArchiveTeam submits the data to Internet Archive, which anyone can upload and download from. This data is continously being uploaded and made public and free. See https://archive.org/details/archiveteam_greader for example.
Anyone can do anything with that data. Your OPML files are not being submitted though. That's also being said on the linked site for this item.
If you feel like you'd want to submit them, but hide.. I don't know, that some feeds might be relevant to each other? or something like that: then you could, split up your list of feeds and submit them in chunks that make sense to you. From different IPs or what not.
All of the historical data from going through all feeds through Google Reader, will be uploaded to the Internet Archive.
That means it will be available to everyone.
This is a non-profit service run by volunteers, that believe in saving data - because there's smart and creative people around the world (High concentration on HN) that can do good things with data.
I can think of one example: All of the new RSS Reader services could slurp this data in and provide you with a better service (and they won't know the feed URLs came from you)
It's not all or nothing. OPML is easy to edit, and I did just that before I uploaded my OPML - deleted anything which might be a security issue. It took like 10 seconds.