Internet is losing its memory: Cerf
itnews.com.au
itnews.com.au
Take Slack. You can install it locally, but that doesn't matter, because it won't start without an internet connection, and its local storage is completely opaque. Compare this to a mail client, where everything is stored locally, indexed and searchable offline, and can be exported into a universal albeit messy MIME format, as well as imported into other accounts.
Of course this is necessary for the business model, the Slack free-plan event horizon wouldn't be effective if it only applied to which new data you can sync down... and if users only discovered this when e.g. moving to a new computer, they would quickly start to wonder why they can't just transfer the data they already have themselves.
With Google Drive, your 'docs' are just placeholders pointing to the cloud. You can export them individually to Word or print them to PDF, but it's a manual process, which you will likely only think of when it's too late.
For an example of how this can really matter: an acquaintance is embroiled in a legal dispute with an ex-employer. Thanks to Apple's sane implementation of IMAP Mail on iOS, she still has access to all her company communications there to use as evidence. Unlike on her PC, where she was just using webmail the entire time and has nothing.
The incentive in the cloud age is to create dysfunctional products that provide an illusion of permanence instead of the reality of tangibility. I expect this is only going to become a bigger problem over time.
Actually it's really really easy to do in bulk: https://takeout.google.com/settings/takeout?pli=1
Not invalidating the rest of what you said.
I'm desperately trying to connect to "modern" chat services via things like Pidgin to have logs - while on some level, remembering everything exactly the way it was is unnatural, links and knowledge is important to be possible to be kept.
I think people are starting to realize that actually having a copy is important. Hosting things for yourself, taking care of backups, etc, are painful, hard problems, but they worth it in the long run.
One more thing: many digital media is significantly more ephemeral, than a lot of us realizes. Out of countless CD from the past 20 years, I can only still read a few of them (the ones written with a 1x Plextor SCSI drive are all still fine). If you truly value an photo, make a proper print, archival grade tint, archival grade paper.
Or you could parity-pad your backups so even 20% corruption is still recoverable and periodically renew them. Several orders of magnitude cheaper.
Isn't this what SaaS is all about?
It's a tired tune to say that RMS was right, but he was. Letting people be controlled by software is not a good idea for pragmatic reasons that go beyond mere morality.
Also, I would not equivocate not controlling your software with being controlled by it. In some cases maybe, but definitely not all.
Anyway.
> With Google Drive, your 'docs' are just placeholders pointing to the cloud. You can export them individually to Word or print them to PDF, but it's a manual process, which you will likely only think of when it's too late.
This has been trod to death over the years. Let's do some find-and-replace.
With a Hard Drive, your 'docs' are just placeholders pointing to the platter. You can copy them individually to diskette or print them to hardcopy, but it's a manual process, which you will likely only think of when it's too late.
I agree this is an issue, but just thought it was worth pointing out that (cloud + offline) is better than just offline.
I'm a bit of a digital packrat. Despite carefully preserving and porting my data from system to system, I lost years of email archives in the aughts because Outlook Express would prompt me to archive old mail, but I didn't realize it was over-writing the old archive until I went looking for a message and couldn't find it (when it was too late). The lost email covered all but the last few months of my undergraduate degree.
With the right spec, you can have Google Docs style real-time collaboration or Slack style chat while allowing users to own and archive their data, remaining resilient under arbitrary network conditions and topologies, and retaining full edit history with the option to link to any previous revisions or even individual changes (with context). The system could be built on top of databases, files, or both. Your "*.crdt" documents could live in your Dropbox and work seamlessly with any software that understands the spec.
It'll take some work to get there, though. And of course, it'll never be as resource-efficient as a centralized architecture. But today’s devices can more than handle the strain.
As I've mentioned in another HN thread, I've been using youtube-dl to back up YouTube videos that are important to me since it's increasingly obvious that YouTube is merely a profit machine that would delete all of its user's videos in one fell swoop if they made an extra buck off it. I've also begun to back up some blogs that have vast mounts of great content, but I've found that HTTrack is... well, terrible. That's not to say it doesn't technically do its job, but it's way too aggressive. I wan't something that's easier to configure and can do a better job of taking "snapshots" of blog entries while blocking ads, javascripts, being able to resume properly if my power goes out, etc.
When you say personal archives, I wonder a little about what we'd lose to filtering by people who know they're building an archive for posterity, versus people in the future rediscovering things that were safely forgotten for some long span of time.
At the delta of personal archives and projects like this, I imagine something a little like SETI or folding for home archival (is this just describing IPFS?), but I think one of the challenges is how we preserve the right kind of data. I did some spring-data-cleaning this year while building a new system, and one of my tasks was addressing duplicate files. This forced me to confront hundreds of gigabytes of VM images, snapshots, and massive piles of duplicate files from various installations of Ruby and Python.
But perhaps I underestimate the value of even the duplicate VM images or language installs. Perhaps, centuries hence, researchers will hopefully sift yottabyte of data looking for all of the dependencies to successfully resurrect applications they need to experience in order to understand references to them in well-stewarded tweets, posts, or articles.
And then there is the copyright issues where archiving is getting more difficult. Kids today don't know what a great tool google cache was because it's gone. And archiving sites are being attack from news, media, politicians, etc.
In other words, the internet dark age isn't going to be a result of formats getting old ( though that is an issue ). It's going to be a result of us only being allowed to see through the frame that a handful of companies deem appropriate. The dark ages didn't happen because formats got old. The dark ages happened because the powerful decided that we should view the world within a certain narrow religious frame and censored everything else.
Browsing Internet Directories.
Just plain finding some guy's massive meticulously-maintained HTML-only site full of fascinating articles, bookmarking it, slowly working your way through it over the course of your evening browsing sessions.
It also does its best to make that difficult. Advertisers luring people away from what they were doing. Content marketers spamming the web to the point that a regular person is more likely to be trapped in some bullshit low-quality clickbait article than to get correct information. You need to be hyper-focused just to get what you're looking for.
Only papers from famous people were kept. Even so, for example, HP's historical archive was all on paper and placed in a single building. Which burned down.
The WTC collapse destroyed the unpublished archive of Kennedy pictures.
> a 2013 survey of law- and policy-related publications found that, at the end of six years, nearly fifty per cent of the URLs cited in those publications no longer worked. According to a 2014 study conducted at Harvard Law School, "more than 70% of the URLs within the Harvard Law Review and other journals, and 50% of the URLs within United States Supreme Court opinions, do not link to the originally cited information." [1]
I doubt that the US Supreme Court was regularly losing track of cited documents before "information technology" came along. A lot of documents may be recoverable even if the original URLs are broken, but is this really the best we can do to manage information in the information age?
Then again material published in journals and books is greater than material published on the internet, if for nothing other than it can be relied upon to be found in decades to come.
Although it is interesting (and a bit depressing) that we can't really seem to escape that model for preserving knowledge over the long term. Physical media needs patronage, buildings, printing presses, etc, while digital media needs infrastructure, manufacturing, programmers, etc.
At the time of Gutenberg, there were about 30,000 books in all of Europe. (Inthink -- research suggests this may be higher, see mss. chart below.) Not titles, but books: individual, discreet, volumes. The University of Paris in 1200 had on the order of 2,000 volumes, amongst the largest collections of the time.
There were something shy of one billion volumes by 1800. Publishing in England during the 19th century was about 1,000 titles/yr. For much of the 2nd half of the 20th century, the US Library of Congress added about 300,000 titles/yr. That's remained fairlyconstant, though "nontraditional" publishing (self-published and on-demand titles) take this to over one million titles annually.
https://upload.wikimedia.org/wikipedia/commons/thumb/e/e4/Eu...
https://upload.wikimedia.org/wikipedia/commons/thumb/2/24/Eu...
Google has counted the total number of books now in existence (titles), and came up with 129,864,880.
https://www.telegraph.co.uk/technology/google/7930273/Google...
http://booksearch.blogspot.com/2010/08/books-of-world-stand-...
The Thesaurus Linguae Graecae is a comprehensive archive of all known surviving Ancient Greek literature:
Today the Online TLG contains more than 110 million words from over 10,000 works associated with 4,000 authors and is constantly updated and improved with new features and texts.
http://stephanus.tlg.uci.edu/tlg.php
(Many of the authors survive only in fragmentary quotation.)
Given that 100k words is a substantial book (about 400 pages), the entire surviving Greek bibliography -- every retained written word -- would be 1,000 such volumes.
A bit under 5 GB of uncompressed text.
The chemical stability of clay means that, yes, some records have survived. They are a minute fraction of all ever created, and those represent a minute fraction of all information that existed to be recorded, but never was.
The largest surviving cuneiform archive is at the British Museum, comprised of about 130,000 tablets, though the core collection is about 30,000. Each tablet is more analogous to a page than a book.
https://members.bib-arch.org/biblical-archaeology-review/31/...
https://www.bl.uk/aboutus/legaldeposit/introduction/
"By law, a copy of every UK print publication must be given to the British Library by its publishers, and to five other major libraries that request it. This system is called legal deposit and has been a part of English law since 1662.
From 6 April 2013, legal deposit also covers material published digitally and online"
Its not crazy to think that, if we live that long, we will figure out a way to escape even the end of the universe (I'm reminded of Asimov's The Last Question). But that progress definitely won't happen if future generations can't build on what we've learned.
I actually think living with the idea that everything is temporary is very very dangerous. Nothing is temporary, except you. We need to live with the humility to recognize that our life on this planet is incomprehensibly short, and we need to spend it doing everything we can to give our children the headstart they need to do everything they can to give their children a headstart... into infinity. That's the only way we improve as a species. That's why there is no greater sin than destroying or limiting access to information, and conversely why the internet is literally the most substantial development in the history of mankind.
The Last Question is a work of fiction. All works of fact seem to support the concept that it is not possible. That isn't to say that we will never discover properties of the universe that open new possibilities, but A) I think eternity is just a different kind of ending anyway, and B) stubbornly grasping to such a concept seems a desperate attempt to deny one's own mortality.
I'm not advocating for forgetting things for the sake of it, I'm advocating against the idea that everything must be preserved. Certainly we should try to keep in mind the lessons of our past, but we should also be willing to let go of things that haven't given us good reason to keep them around.
Accepting mortality is important for a lot of reasons I find difficult to explain properly. It's a matter of shared context being the basis for communication and the requirement for a whole lot of concepts to have to be understood by both of us, and referenced by the same words, to convey actual understanding. It is difficult, for instance, to separate the topic of mortality from the discussion of the conception of self. I will attempt an explaination regardless.
Knowing, and accepting that your end is inevitable is important context for how you choose to interact with the world. Let's try a thought experiment: There's a version of the world you want to live in, and there's alternatively a version in which you don't, right? Since you will someday die anyway, it makes no sense to survive at the cost of moving the world towards the version that you don't want to live in. If you don't accept that the end is inevitable, you can justify such acts anyway because you'll always be able to move the world the other direction later, just as soon as the current existential threat is dealt with. Only, of course, there are always more existential threats, whether real or imagined.
Many atrocities were and are committed in the name of survival, whether of an individual's physical self, or one of their shared memetic selves like culture or society, their genetic self embodied in their children, or sometimes even more raw memetic concepts. That's the "rational" way of interacting with the world. You have a goal (survival of some conception of self), and you do anything to achieve that goal, because rational approaches do not accept failure as an outcome. Accepting mortality is recognizing that the goal is ultimately unachievable, and because of that, there are potentially things of greater value than survival. You don't have to survive, failure to achieve that goal is an acceptable outcome, and you can choose to interact with the world differently.
You mention physics. A constant in our universe is that things decay unless actively maintained. All life is fight against entropy. Trying to preserve things is only natural.
We put priority on benefits we can realize. Short term benefits are easier to realize, and decisions made in the moment, absent malicious intent, make sense in the moment. Hindsight, as they say, is 20/20. We make a prediction now what will be important later, sometimes we are right, sometimes we are wrong. Perfection is impossible, and even if it weren't, it would not be worth the cost of trying to attain it.
That said on the topic of conservation, or more accurately History, I don't think you can convincingly argue for forgetting things as a society. The cost of repeating errors are as great as the benefits of safekeeping knowledge.
And how might you go about systematically identifying them?
And what might your incentives be for doing so?
— Douglas Engelbart
If an advanced alien civilization came to Earth, and helped ancient human civilizations to built and operate a computer system, which can share, store and translate scientific discoveries all across the globe (while they are not allowed to help humans besides operating the system) throughout 5000 years of history - how could it change the human history?
https://adactio.com/journal/11937
"The original URL for this prediction (www.longbets.org/601) will no longer be available in eleven years."
This has been a known problem for a while, but it's always nice to see more voices crying out about it to the more general public.
Also archive.org is sometimes on the brink of legal/illegal, e.g. not all material that can be found on archive.org is really legal; though for lots of the formally illegal archived content, the copyright holders do not care or do not want to cause an outcry.
https://bugzilla.mozilla.org/describecomponents.cgi?product=...
Flash is fading from mainstream use on the web (thankfully), but not hard to run Flash today if you want to.
I’d never heard of the Digital Object Architecture (DOA). The article itself makes light of the unfortunate acryonym with “History pronounced DOA”. That actually left me confused about what they were talking about for a minute.
No idea if it’s a good way to preserve academic papers on the internet, but the business model side of it is a pure open question. That makes me wonder whether it solves anything at all. The problem with information on the internet is that the people who publish eventually lose the interest or the ability to continue paying for storage and access.
“Economics/Business Model
While the Handle System has been used for many years in publishing and library systems, generalizing to other applications, e.g., Internet of Things, will likely generate economic concerns related to the business model of the system, especially at the Global Handle Registry. Will organizations be charged for each identifier? Will organizations that acquire a prefix be able to create unlimited sub-prefixes or will they be charged for each sub-prefix? How will these policies be developed? How will the money flow? What will be the impact on developing countries or small businesses?”
https://www.internetsociety.org/resources/doc/2016/overview-...
Discussing about the 4-color theorem recently proved, latest version of C compiler, difference between a vacuum tube amplifier and a solid-state amplifier, where GNU Project and Linux kernel was launched, and early online culture and tons of colorful, hilarious, but forgotten and buried memes, and weird phenomena emerged from the collective (un)consensus... Sci-Fi fandom being an integral part of online and hacker culture, millions of lewd story written in alt.sex, "Immediate Death of Usenet Predicted!", "There is no cabal", alt.french.captain.borg.borg.borg, Coffee and Cat warning, The church of Kibology, Anti-spam Movement, creationism vs evolutionism debates at talk.origin, Meow Wars - the first meme war online, all the personal attacks, trolls, flame wars, and "cyber-stalking", etc.
Then centralized WWW replaced distributed Usenet, crappy HTML replaced perfect machine-readable data format. Would we have a similar archive for Reddit or Hacker News? Possibly not. So Hacker News, just come and create one! You can make it! Anther unique challenge created by WWW is the inaccessibility of server-side software - exporting and preserving the data is NOT enough, unlike Usenet which you can just load any data. The user-interface and functionality of one website itself is also the collective memory that needs to be preserved - we need replicated software of a website, which has identical user-interface, which has all the functions from the original website: users to click an username and see the posts, karma of this user, etc. I don't think anyone even noticed the existence of this problem. Luckily, major websites online such as Reddit or 4chan, all use FLOSS software which would make the work easier, but still a huge challenge due to the inaccessibility of raw database. Also, to make some contents meaningful in the future, external resources such as hyperlinks to other websites and images should also be preserved, considering this, the chance of creating an authentic and complete archive is even lower.
---
But even if we're still using a distributed network where data preservation is still technically possible, and there is no walled garden, it may still be difficult to implement. In the era of Usenet, you often attach your name, address and phone number - there was virtually no threats except for a few trolls - this is why archiving Usenet was possible in the first place. But the Internet is not the Net anymore, now not only humans - almost every piece of equipment involved on the route may be your enemy.
The ongoing security and privacy movement is a huge threat of historical records. From my observation, at least of infosec hackers community - After Snowden's revelation, public and open discussion is slowing being transformed into private, closed, encrypted and temporary activity, plus self-hosted platforms like ActivityPub, GNU/Social, Mastodon. This is indeed good from a security and privacy perspective and it is exactly what we need now.
But we are also creating a huge gap of knowledge, information and history on the Internet. After my death, none of my self-hosted code, or my blog, or my GNU/Social posts will survive. In conclusion, "collect 'em all" is both an malicious NSA dragnet surveillance, and a glorified act of history preservation. This is where the contradiction lies.
I don't know what to do. For WWW, archive.org is a workaround and I think it needs more donation. But for all the other self-hosted things like Mastodon and git server, there is no solution at all.
See also: https://en.wikipedia.org/wiki/Newsreader_(Usenet)
Note that olduse.net is a playback of Usenet, time-shifted 20 years back, some of the events I mentioned has yet to occurs, you can navigate the website, download the original archives, and load it by your own to explore.
If you don't have the setup yet, to get an quick idea of how it works, read these interesting articles.
* http://olduse.net/blog/what_rms_saw/
* http://olduse.net/blog/Dennis_Ritchie/
* http://olduse.net/blog/stargate_controversy/
You can also just browse the old Usenet from Google Groups. It's the same contents anyway, but the experience is poor.
I’m all for making it easier to recover data and generally make it easier to store stuff. And I believe you have a right to your data (yay GDPR). But this isn’t a catastrophe.
That's an oft-repeated but ultimately meaningless statement.
We know how to make things better than Damascus steel or Roman concrete. Our processes have exceeded the ancients. The quality of Damascus steel is likely overrated, since it wasn't quite as absolute shit as what was being regularly traded at the time. Same with concrete.
We might not have exquisite written recipes and procedures for these materials, but they have been reverse-engineered, and it turns out, they weren't that great compared to modern chemical and material engineered products.
We don't need to 'rediscover' Damascus steel because it is obselete. A romantic idea, and poetic in how it was 'lost' to time, but consider that it was 'lost' to time, similar to Japanese steel-working, because nobody wanted to buy it anymore. It became economically and culturally irrelevant.
And that will be true of essentially everything we forget in the future as well — the forgetting is the sign that it wasn’t needed.
Looking forward to the next few years and their (anticipated) growth