Archive.is owner on “continuity of his project”
blog.archive.today
blog.archive.today
I'm reminded of some non-profit organization that was forced to shut down their websites because they ran out of money. In retrospect they could've setup a trust fund early on, stuck all their money in there, and then had a perpetual annual income from that for operating costs, instead of spending down. All fundraising could've gone into the trust fund in order to boost the annual budget, etc.
>How much does hosting cost you per month at the moment?
>about ~$2600/mo of pure expenses on servers/domains, not counting “work time”, “buying laptop/furniture”, etc. ($100…300/mo covered by donations + $300…500 by ads)
https://blog.archive.today/post/659383959382294528/you-said-...
There's a huge difference between "please donate to our project" and "it costs us $X/month to run it, we have Y users and managed to collect $Z so far, please donate".
The first one will get you about $1 per 10K-1M users. The second one will get your goal fulfilled, as long as you are reasonable, and have enough users. All it takes is a noticeable message and a way to update it automatically based on the money received.
If the concern is perpetual access to archived content but under your terms, that is where the cost comes in. Somebody somewhere is paying for power, cooling, connectivity, and disks. The Internet Archive estimates it costs them $2/GB to host data uploaded in perpetuity. Please consider donating if you're uploading content for permanent archival and/or deriving value from hosted content.
ipfs's economically incentivised hosting layer (filecoin) offers storage in monthly increments, not perpetual.
practically yes, but
A system of mirrors prevents a single node going down from taking the whole system (in theory at least, we've all seen plenty of times where failover goes poorly), but it doesn't do anything to ensure the long term survival of the system as people lose interest, lose the ability to participate, and sometimes die.
If you were around certain internet forums in the late '00s you might have run in to an image hosting platform called WaffleImages which was created in response to yet another popular free image hosting service locking down their embedding and ruining thousands of old posts. The goal was to distribute image hosting among community-operated mirrors, and it worked great for a few years. Over time though people lost interest while the rate of new mirrors getting added dropped to basically zero and eventually it fell apart.
Cool URIs don't change: https://www.w3.org/Provider/Style/URI.html
You can believe that cool URIs don't change or you could go the IPFS route. Similar to the way torrents have a 'health' score of plenty of seeders, and IPFS resources could live as long as people want that resource to exist (Not sure if that situation is baked into IPFS though).
Also: have you looked into Filecoin?[0]
I love the site, but his stance on this doesn't really make sense to me, and it's a shame that millions and millions of people use 1.1.1.1 daily and archive.is is the one website that doesn't work for those people.
As for why archive.is cares so much...that I don't know. Perhaps they rely on such data to give a fast experience, and are tired of this charade...but that's just speculation.
There's nothing incorrect about what Cloudflare is doing, EDNS does not require ECS data to be included in requests, but for whatever reason the maintainer of Archive.is decided to block 1.1.1.1 over it.
edit: from here https://news.ycombinator.com/item?id=19828702 I gather that this indeed harms CDNs outside the ones that Cloudflare has a business relationship with.
> EDNS IP subsets can be used to better geolocate responses for services that use DNS-based load balancing. However, 1.1.1.1 is delivered across Cloudflare’s entire network that today spans 180 cities. We publish the geolocation information of the IPs that we query from. That allows any network with less density than we have to properly return DNS-targeted results. For a relatively small operator like archive.is, there would be no loss in geo load balancing fidelity relying on the location of the Cloudflare PoP in lieu of EDNS IP subnets.
> We are working with the small number of networks with a higher network/ISP density than Cloudflare (e.g., Netflix, Facebook, Google/YouTube) to come up with an EDNS IP Subnet alternative that gets them the information they need for geolocation targeting without risking user privacy and security. Those conversations have been productive and are ongoing. If archive.is has suggestions along these lines, we’d be happy to consider them.
would not be surprised if he has some personal axe to grind with cf (they are no sheep either).
also i would be wary of overestimating market penetration of any 3rd-party dns provider; iirc google has total dominance of this segment and is still below 10%.
All-around a huge accessibility impediment.
They are the next Google in terms of "evil companies"
archive.is was also a great mirror to instagram and linkedin for public profiles, but it doesn't archive instagram anymore.
"There is no Instagram content which don’t need to login.
If you can access the page without login, it is sort of “promo preview“, after few pages accessed this way, they add your IP into “promo is over“ list and will redirect to /login on every future request.
I just have not enough fresh IPs to abuse this mechanism."
https://blog.archive.today/post/659927354404192256/instagram...
Archiving everything is still a novel and not fully understood concept it is not that clear that it is useful or beneficial over long term.
Forgetting can be a bliss.
I've helped with a couple risk registers at tech companies. Two things I've never seen appear in a risk register: The company runs out of money. Human society is wiped out. I've been laughed at once for bringing up variations on these. They're out of scope; risks stop being a threat when there's no one left to care about them.
I think the goal of keeping an internet-scale level of data accessible and searchable, for longer than one lifetime, is an impossible task. Maybe Archive.org/Archive.is can pull it off; I doubt it. Its an insane amount of data. Most of it is totally pointless, but its really difficult to pick-apart what's useful and what's useless, so you have to keep as much as possible without bias. All of that is on hard disks which violently spin around at 8 meters per second, accessed by software which we all know breaks every day but are too afraid to admit it, over a network of other computers with all the same flaws, distributed globally, yet can be significantly disrupted by one roadside construction worker and a jackhammer.
The internet didn't increase the lifetime of data; it decreased it. Sure, we have far more of it at our fingertips than any other point in history, but that's not lifetime; that's just volume. And that volume has desensitized us; its fundamentally impacting our innate biological memory capacity, and the social structures we form around memory. We know the Library of Alexandria existed because people wrote about it; the pages laid for thousands of years; its memory passed verbally from person to person.
If all computers stopped functioning tomorrow, not even disappear, they're still there, they just don't work: Would the memory of Stranger Things still be known in two thousand years? I doubt it, but: if the only thing which offers us a satisfying "Yes" is "we keep the computers running, accessible, indexable, searchable"; that seems, at the very least, given the extreme challenges we as a species will be facing over the next century, beyond the scope of human possibility
you are stating the current status quo as unavoidable - why is that?
Archiving everything has a massive cost and if it were illegal to archive unless people consent it would be much harder to get away with it.
When I post to a public forum I gave that forum rights to publish what I said, I did not give everyone else the right to store what I said
The crux may be whether websites that require you to login to use them, like Facebook, are considered public or not. But anything you say or do in Facebook that you don’t restrict to being viewable only by 1st degree friends is probably considered public.
That'll be -5 points for questioning the social credit system, Citizen.
I'm not sure about consent, but presume it will be 'stuck' and un-removeable from the net once it's out there. (So be careful what you disseminate). Some people even go out of their way to make sure certain content will never be forgotten from the web.
There is no practical reason why a forum could not have an optin-optout-delete API/protocol/etc that all search engines/archivers should follow.
When I delete something all should follow the orders and those that do not need to be responsible for it.
Of course there are plenty of business reasons why engines don't want to do it.
When you write a brief or a book, it's a performative act; you cannot pretend nothing was said.
The Internet is textual media, just like books, and unlike television or talking in the park.
If you don't like it, you can only burn books or drive modern technology in that direction
The reason is that bits are not physical. Anyone can copy them and re-upload them, without any cost.
Furthermore the tide will be ordered back, the contents returned to Pandora's box, and universal entropy decreased. Failure to comply will result in a fine.
There are all kinds of feedback cycles going on. Neither of these two are rational:
1. There are no ways to correct the problem you describe
2. The problem will correct itself without individuals and institutions making it happen
https://www.quora.com/How-long-does-it-take-for-paper-to-dec...
I use Gmail right now, but I could move my domain's email to Fastmail or another provider without _too_ much work.
Of course, it's also a good idea to back up your email archives too, since that _can_ go away with the loss of a provider's service.
All things are ephemeral after a certain point but archiving typically lasts much much longer than the human operators. Likewise documenting the process and barriers to overcome will help people in the future solve the problem (and a broader amount of people).
This doesn't have to be public, just needs a way to become public.
The blog is at blog.archive.today, it calls itself the "archive.is blog", but when I visit archive.is or archive.today, I'm brought to archive.vn. When I click the "archive.today" logo in the header, I'm taken to archive.ph
Site owner said once archive.today is the domain to use when linking because it will automatically redirect to the correct one.