Public appeal to backup Climate Monitoring and Diagnostics Laboratory's FTP repo
reddit.com
reddit.com
Where's a good place to put data like this? Is there even a place where people can start to inventory what's at risk?
The biggest problem is not knowing what's out there, what records need to be protected, as there's so many people doing important work.
Archive Team keeps a constant eye out for major content sites that are at risk of being shut down and proactively mirroring their content. http://www.archiveteam.org/index.php?title=Main_Page
On EPA data https://twitter.com/textfiles/status/824110893034311680
They don't just archive online content either. They spend many resources scanning magazines, ingesting (and cataloging) CDs and floppies before they all rot away.
You know when the Internet Archive itself is making a contingency plan to move out of the US that things are truly fucked.
The most difficult thing here is any prirotization they end up doing must be kept absolutely secret or you're going to be providing a hit-list to the administration should they ever want one.
[1]: Seems like that happened here: https://en.wikipedia.org/wiki/Wayback_Machine#Netbula_LLC_v.....
Problem is, those use cases are greatly outweighed by ones where either:
1. The site has changed ownership to a different (legitimate) company or organisation, and they don't want their new site archived.
2. A domain squatter/seller has bought the domain, blocked all bots to stop the holding page being indexed in Google and accidentally blocked the archive in the process.
3. A technician or developer has accidentally blocked the archive/all robots due to a personal mistake/copying code from the internet.
4. The site owner doesn't know the archive has a bot, and has blocked all non Google/Bing bots to 'reduce strain' on the server. That last one is depressingly common:
http://webmasters.stackexchange.com/questions/75993/only-all...
https://twitter.com/internetarchive/status/82406378479265792...
So in this case, the Internet Archive should work fine.
Otherwise, I guess AWS Glacier if you want to pay for it.
It needs to be on hardware an organization like the Internet Archive controls or at least ordains. We're talking bunkers in Sweden or Iceland or some country with the political backbone to defend this sort of scientific data.
*Not sure if this is the right term.
From a current session:
ftp> open aftp.cmdl.noaa.gov
Connected to aftp.cmdl.noaa.gov.
421 There are too many connected users, please try later.
ftp>
Seems like archive.org should be archiving this stuff, so it's all in one place and future researchers don't have to go hunting all over the internet for it.
- In the US, the CFAA has language that may or may not be relevant [1]. This isn't intended to be FUD, although I can understand if you think so. As with most legal things, don't believe me -- a random guy -- on the internet; ask a lawyer, or at the very least be aware of your risks.
- Check license and terms, if included. Reddit post claims these [2] are terms, seems okay, but I have not confirmed this for myself.
- FTP is unencrypted and non-tamperproof, so data integrity and data authenticity, and connection privacy isn't guaranteed. And because the original source did not publish checksums over a secure channel, there is no way to know whether these files are the originals. It will be difficult to prove the provenance and accuracy of the data, despite having lots of third-party copies, without the original source having safeguarded the strong cryptographic checksums the entire time, or having published them over a tamperproof channel.
[1] https://www.law.cornell.edu/uscode/text/18/1030 [2] https://www.reddit.com/r/DataHoarder/comments/5q4xxe/erik_fi...
They needn't even be strong, right? Even a weak checksum on a data set of this size is basically collision-proof, no?
Any other function, whether non-cryptographic, too short, broken or having known weaknesses, while perhaps suitable to detect accidental corruption, will not protect against the former.
For example, SHA-384, or SHA-3 are great choices.
Science is the enemy. It will be destroyed if it continues to get in the way of the administration.
It happened here (Canada) http://www.macleans.ca/news/canada/vanishing-canada-why-were... and it could happen on an even bigger scale in America. The Conservative party went about cutting funding, slashing archives, burning everything to the ground, sometimes literally. They'd do it with little warning, zero fanfare, and an aggressive timeline. One day you had a climate archive, the next the shredding company had taken care of it.
Those servers aren't free. They depend on budgets and grants to stay running. If that money is cut, those files are gone.
This isn't hysteria. This is insurance. Don't be That Guy.
It isn't alarmism to have a contingency plan if everything goes to hell. Much like Peter Thiel has one.