Thank you for helping us increase our bandwidth
blog.archive.org
blog.archive.org
I donate to them monthly and know a lot of other people do as well, so I don't worry much about their financial stability. I'm more worried about external pressures taking content down. I hope the data is backed up six ways to sunday, and that somewhere there's a plan to make it all accessible if Internet Archive can't continue to play the role it does.
Yes I save some dead sites that I think are critically important resources to tarballs that I'm storing separately in case the Wayback Machine disappears.
Your stance is pretty extreme - please don't discourage people from making backups.
im3w1l, HN, your keyboard driver, CPU, firmware, motherboard, ISP, Google
One of these things is not like the others.
As far as I've ever been able to determine from talking to anyone at IA (e.g., Kahle, Scott), they don't really have any sort of backups that could actually be restored from in a disaster situation.
Furthermore it would not be so hard to translate Archive.org items to IPFS objects, if there were an effort to pin a significant number of them to storage and network.
(50 petabytes * 0.2% = 100 terabytes)
[$0.005 ($/GB/Month) BackBlaze cost]
[(50 petabytes) / (1 gigabyte) = 50,000,000]
(50,000,000 * $0.005 = $250,000 US$)
—————
Meaning based on my numbers, that is $250,000 USD a month to host 50 petabytes of data on BackBlaze.
Also, those backups would be (relatively) cheap to keep, but not necessarily to restore.
I had constant issues with objects simply never (hours and many requests) being found, despite being pinned in several places; and the daemon sucked resources away from the system at an alarming rate back when I was trying properly.
With all that being said, maybe now is the time to look at it properly, the budgets are there. Maybe Juan, _prometheus, can find somebody to at least PoC this important application.
IABAK stores 100 TB, or 0.2% of it.
I also remember reading about Sia on HN, which is a dapp that pays hosts to store data and distributes it. Looking at the going rates on Sia ($1.45/TB/mo), that's $870k/yr. That's ~10% of the IA budget (which is only $10MM/yr, which sounds very efficient!) but shows that the order of magnitude is not that crazy.
Azure Archive is $2/TB/month ($1.68 if reserved)
AWS Glacier Deep Archive is $1/TB/month
GCP Cloud Storage Archive is $1.20/TB/month
Of course, there can be i/o and network charges, and different levels of redundancy (but possibly bulk discounts)...but the bare storage costs for for 50 PB per year would be roughly $600k - $3 MM/y.
https://docs.aws.amazon.com/snowball/latest/ug/create-export...
is that a "Million Million" eg. 10^12 ?
In this case, 6250*150 = 937500, almost a million.
Surely a lot of people would be happy to donate unused space on their drives (I know I would) - especially since backup wouldn’t use much bandwidth...
https://blog.archive.org/2016/12/03/faqs-about-the-internet-...
I personally think it’s great but surely companies aren’t too pleased? How does archive.org avoid being sued into oblivion?
Which seems smart enough: minimize litigation costs while not losing content permanently via court orders or dmcas or similar threats.
But it's frustrating for the wayback machine- where sites like Snopes and some newspapers have opted out of having their history published (after being accused of ghost edits to articles).
I don't know what the right answer is, but even if they didn't display the page, it would be nice to see the diffs (ala wikipedia), or at least show if and when changes to a page were made. Maybe that's beyond their mission scope, though.
Is there any way to interpret this behavior as anything other than straight up admitting to being dishonest? Especially in context of accusations that could be easily disproved using the Wayback Machine... unless there is something to those accusations, that is.
> Computer programs and video games distributed in formats that have become obsolete and which require the original media or hardware as a condition of access. A format shall be considered obsolete if the machine or system necessary to render perceptible a work stored in that format is no longer manufactured or is no longer reasonably available in the commercial marketplace.
But here is a complete NES ROM set, just sitting there, on archive.org: https://archive.org/details/NESrompack
I just find it strange. Nintendo is vehement about their IP. For archive.org to get an exception doesn't make sense to me.
I don't know how archive.org gets away with that.
But it probably needs tweaking for that purpose; for one you need to ensure that the data is evenly distributed, and second you're dealing with data that is appended to regularly.
But I think it can be done.
Here's an overview: https://www.archiveteam.org/index.php?title=INTERNETARCHIVE....
Sadly, it seems the project has fallen out of repair.
It's the same sort of feeling as I why I enjoy this site, wikipedia, and so on. Bonus, IA is also a bit of an organizational mess, but rewards the adventurer with rich treasure.
(I was recently on a Korean history kick and came across not only 1, but several entirely different first hand books written by visitors in the late 19th and early 20th century, scanned in, freely accessible/downloadable, in a variety of formats, and with an excellent on-web reader. These books are so out of print I checked with three local counties for copies and none of them even have references to any of them in their catalog -- treasure!)
[1]: adding up the outbound numbers for each datacenter: 1.888 + 8.003 + 810 + 1.958 + 807: https://grafana.wikimedia.org/d/000000605/datacenter-global-...
[2]: https://en.wikipedia.org/wiki/List_of_Internet_exchange_poin...
[3]: Not 100% correct as thanks to Corona, ix.br now has passed 10 Tbit/s peaks. https://ix.br/noticia/releases/ix-br-reaches-mark-of-10-tb-s...
I should have added: 60 Gbit/s is a lot for a donation run service like archive.org. Not sure if there are donation run services on the internet with more traffic.
Looks like they max out at 100 Gbps, or as low as 40 Gbps, depending on appliance model and link aggregation configuration. No argument either way, just thought it was cool info.
https://openconnect.zendesk.com/hc/en-us/articles/3600345383...
Netflix has a (relatively) small number of large content pieces, and a CDN node would hold the hottest items.
IA has 430 billion individual pieces of data, on average much smaller than a piece of Netflix content.
So each "byte" transferred from IA is "more work" to produce, and less likely to come from some kind of cache.
Both your upstreams, Cogent and Hurricane Electric, offer 100G ports at fiveish grand per month in carrier neutral DCs. Given that your budget is in the millions, an outlay of this magnitude doesn't seem wildly out of the question.
If you can explain what the problems are in getting more bandwidth, I'd be more than happy to see what I can do to help.
Please let me know if you'd prefer to discuss the matter over email.
https://mailman.nanog.org/pipermail/nanog/2020-May/107720.ht...
I hope they'll look into it, find a way to identify that traffic, and then (ideally) classify it accordingly and continue to serve it only after all human traffic has been served (i.e. if they run into a bandwidth crunch due to bots, only the bots suffer, and they don't need to upgrade as quickly).
There's also at least one outright correction that only seems to exist on the Internet Archive now: https://web.archive.org/web/20200428214306/https://www.bbc.c... "Correction 25 April 2020: An earlier version of this article incorrectly said that France had conducted just under 140,000 tests a day by 21 April. The figure of just under 140,000 refers instead to the number of tests it had carried out weekly."
("Journalist" is in quotes because, while that was her job title, she specialized in abuse, libel, smears, and calls to murder, and I'm glad she lost the libel case)
I do like this a lot for high profile accounts. It seems fair that if you're verified and have a very large number of followers, you are operating at a privileged level of impact and should lose your right to delete tweets any tweets.
I have a side-hobby of setting up scrapers which pull scraped data into a git repository, precisely for this kind of thing. I'd be happy to set a few up.
Some of my posts about this technique (which I call "git scraping"): https://simonwillison.net/tags/gitscraping/
(1) US coronavirus data in general,
(2) examples of sources that do not log prior days data, or
(3) source that “edit” prior data without noting the edits.
(4) something else
If (1) this page in the table under the column “sources” links to where the data came from:
I've also been running bash scripts every day to save an archive of some of those pages https://github.com/jeromegv/covid_data
Also, I'd trust these guys not to mishandle .org
(yeah, I know it's not going to happen)
As much as I like the idea of running a TLD, if anyone gives me a TLD, I'm gonna put an MX record on it and be jonah@org. That's a promise.
EDIT: To be fair, I would probably be overwhelmed with the usual sense of responsibility and not do this. But, the temptation
So one night, a new ham is operating his friend's fancier station, which has the HF equipment to talk clear around the world. And they make contact with another ham identifying himself as JY1. No suffix. Well that's weird, so they look it up, the JY prefix is Jordan, okay. So the guy on the mic just asks, "why doesn't your callsign have any letters after the digit?" and the reply comes back with a chuckle, "Oh, because I am the King."
That was the late King Hussein. about whom much has been written. I don't know if his call ever caused trouble with software validators, but I'd certainly believe it.
Took a while to get ARRL Logbook of the World credit for that one.
PS - Don't forget to pick an Amazon Smile charity and use the Smile.Amazon sub-domain. Donated almost $20 to the EFF that way.
It's also essential for non-historians. A huge percentage of the info people need just to do their jobs or pursue their hobbies is only found on people's personal webpages, and those presumably all go away once the person dies.
They go from archiving the web (simple websites), itself a huge feat, to archiving a bunch of videos from justin.tv as long as they have 10 views?
There's more cases of websites shutting down, and AI feeling the need to archive everything.
Do we really need it?
Then put the file in a torrent. Let the users seed it.
Users can use sqltorrent (https://github.com/bittorrent/sqltorrent) to query the db without downloading the entire torrent - essentially it knows to download only the pieces of the torrent to satisfy the query.
Every time a new dump is published by internet archive, the peers can change to the new torrent and reuse the pieces they already have - since SQLite is indexed in an optimal way to reduce file changes (and hence piece changes) when the data is updated.
I talk a bit about it here: https://medium.com/@lmatteis/torrentnet-bd4f6dab15e4
Would save Internet Archive lots of bandwidth and hassle
How does this compare to something like IPFS?
Would be interesting to learn how it's currently partitioned. I would mimic the same portioning system but use torrent instead so users can help with hosting. And use sqltorrent to serve queries efficiently.
If I was a better software dev, I'd try to make a daemon which I can feed a list of websites which I want to support with my caches that can be upgraded (like the DNS system?). Like a folding at home for bandwidth.
This is almost always the best and most efficient way you can help a cause. You donating infrastructure means donating them extra work - they'll need to integrate and keep track of it, as well as manage a relationship with you (in particular, a risk of problems with you or your service). Meanwhile, with extra cash, they can buy what they need and what integrates well, or pay the most effective specialists they need.
It's an universal principle. Cash is the best gift, if you're gifting to help. That's why e.g. Red Cross frowns at people donating stuff - it creates a huge logistical problem for them as well as depriving them of opportunity to boost the markets in the disaster-struck area. Or why it's better for you to donate money to your local homeless shelter rather than volunteer to work in it - if you work your job for extra hours instead, and donate that money, they'll hire workers better skilled for that task.
I had some hopes for Filecoin (i.e. a P2P system with monetary incentives attached to backing up other people's data, from one of the very few groups in the crypto space that don't look like scammers to me), but I haven't heard anything about it in a while.
Ideally, I would like to be able to point at some RSS feed of torrents of relevant archive.org collections (whatever the curators believe popular enough to benefit from peer-to-peer distribution). I could then set a torrent client to donwload and seed everything on this feed.
It will be an additional back up to our copies. We're hosting a DWeb Meetup to share the latest in Filecoin and Storj (two decentralized storage providers we're experimenting with.)
It's Wed (tomorrow) 5/13 at 10 am Pacific, 5 PM UTC https://www.eventbrite.com/e/dweb-meet-up-virtual-decentrali...
The aesthetic intrusiveness of the archive.org header and footer are minimal since I use a text-only browser that has no Javascript engine.
Sometimes I get "This url is not available on the live web or can not be archived." However this happens for only a surprisingly small minority of websites.
Rarely I find that /save is unsuccessful in which case I can still find past copies using something like
curl -o 1.txt "https://web.archive.org/cdx/search/cdx?url=http://www.example.net&fl=timestamp,original" ;
sed -i '/^[12][0-9]* h/!d;/^[12][0-9]* h/{s/^/http:\/\/web.archive.org\/web\//;s/ /\//;s/\r//;}' 1.txt
The limitation with past copies versus /save is that archive.org will not usually crawl past page one on websites with many successive pages, e.g., http://example.com/?page=2, http://example.com/?page=3, etc.Has anyone ever considered mirroring archive.org, or parts of it, to other geographic locations.
Could this be done. Why or why not.
https://wayback.archive.org/web/*/%S
I'd imagine it would be useful for IA to implement some message for scenarios where a page has already been saved within a certain timespan and provide both a link to the already saved version and offer to save again. As this would mitigate mass savings of an identical page that can occur when some popular link is accidentally shared with the /save/ URL instead of the static URL or when it's a popular page that people want to archive.Archive.is displays such a message (to the effect of, 'this page was archived <date>, if it looks outdated click save') and also redirects to the most recent copy.
HAproxy changes the Host header and modifies the URL. I can either use the text-only browser's http-proxy option or I can direct the request to the web.archive.org backend by adding a custom HTTP header to the request. If I am not mistaken, the so-called "modern" browsers do not have built-in capability to add headers.
If they wanted to use Cloudflare, I'm pretty sure they'd jump at the opportunity.
And would still have to download from their source.
Also, since archive.org has so much content, the caching ratio is going to be very bad and kill CDN efficiency while still requiring lots of direct bandwidth. Cheap direct bandwidth in their case looks best(which is what they seem to be doing).
The way they use bandwidth means the most efficent GB/dollar cost although at the cost of poor performance.
Serving up that large of a media library at mid scale just isn't really a great use case for a CDN, those that have to due to transit costs becoming truly enormous (e.g. Netflix) make an enormous investment in hardware at the edge that probably isn't affordable to IA (or necessary at this point).
Forgive me, but what is an AS?
An AS is an entity significant enough to be reasoned about at the BGP level.
http://blog.archive.org/2020/05/11/what-it-means-to-be-a-lib...
I read somewhere that data creation is exceeding storage solutions' pace. Is this true?
What about a mesh of some kind where every person who install an application hosts bits and pieces of random data and serves it to whoever asks for it?
You can get an update on decentralized storage at our DWeb Meetup tomorrow (5/13 at 10 AM pacific, 5 PM UTC)
https://www.eventbrite.com/e/dweb-meet-up-virtual-decentrali...
Not say it's not reasonable, but how do we work around this? Any alternative services that are more "resilient" in this regard?
https://help.archive.org/hc/en-us/articles/360004715251-Arch...
Semi-unrelated, but if you're looking for ways to help and have a spare server, Archive Team [1] is always looking for additional capacity. Although Archive Team != archive.org, they do grabs of at-risk content which (almost always) get uploaded to archive.org. [disclaimer: I help out with various Archive Team projects, the most recent of which was the backup of Yahoo Groups).
> multiple readers can access a digital book simultaneously
with the only caveat being it is borrowed for two weeks but let's face it, most value of a book comes from its first reading -- and what stops you from "borrowing" it again, anyways.
It's been a gigantic disappointment for me to see them do this and not back but try to placate the authors with weasel words and an opt out.
It’s been a gigantic disappointment to see authors respond negatively to this effort (considering this once-in-a-century event), and I will never buy a book again from one of these authors.
As an aside, many SaaS products have given away their product for free due to COVID and widespread forced WFH.
I'm sure a lot of authors would have contributed their work to the effort, if they'd been asked. But they weren't asked. It's difficult to imagine how you'd similarly force SaaS companies to give away their products for free during the pandemic -- lucky for them -- but if you found a way to do it technically, how do you think they'd react?
You put 'livelihood' in scare quotes, but I don't understand what's deserving of mockery. Writers make their living by selling their writing.
I think that is a fair comparison. As a software developer, I would JUMP at the opportunity to do that. Partly because it would be an opportunity to help sustain the world through this time of crisis at no real cost to myself (perhaps some opportunity cost). And partly because it would serve as free advertising.
Honestly they would have gotten huge buy-in from authors if they'd bothered to ask. What bothers me more is the intellectual dishonesty of their FAQ's.
So a huge group of people get to try your software for free. There's a system to try and cut them off at two weeks but it might fail. This is basically free marketing. All sorts of people who may never have tried your software try it. Nearly all of them don't buy it, but some do. It doesn't cost you any money and you're able to help people in need.
There's evidence that, in general, piracy lowers sales, but the specifics are complex and uncertain[1]. Piracy can have positive impacts on the fortunes of the original artist. I think that a crisis like this is exactly the time to try new and experimental ways of making things available.
Yeah a crisis when the stress level of anyone is already much higher is the best time to crank the stress of authors even higher by experimenting with their livelihood without consulting them!
That said, I really have a hard time assailing libraries. Ebooks are sold to libraries at 2-5x the price I can buy the same ebook for as a consumer. Print books are sold at similar prices to consumers and libraries.
Publish one chapter of the book in an anthology, on the web, etc. I've purchased several books after reading one chapter in this manner.
You mean like shareware?