Google's Vint Cerf Warns of 'digital Dark Age'
bbc.co.uk
bbc.co.uk
But I think that dapper Vint "One is glad to be of service" Cerf is referring to how difficult it will be for far future historians to piece together records of daily life in the 21st century, especially from the view of individuals. Think of Da Vinci's notebooks, the personal journals of artic explorers, correspondence letters from great artists and statesman. Now imagine that those things are stored on floppy disks in Office95 .doc format. We'd have a hard time viewing that media today. In 1,000 years it might be impossible. A lot more will be lost to history than what the Internet Archive is currently storing.
I can't think of a great solution. When I watch some historical documentaries, it seems like every other scene is based on a quotation from a letter or an old photograph. We don't print those things out any more. They'll be lost, completely lost, when email services and social networks finally shut down. And even if people download personal backups, most file and physical storage formats have a shelf life measured in decades at the most. Letters can sit around in attics for lifetimes, undisturbed. A file formatted for a Commodore 64's word processor program might as well be written in a lost language and then one way encrypted. I'm not sure what to do about it, but it seems like a damn shame.
It doesn't seem that important of a thing to preserve though. Sure, if we could hear daily conversations of people from the year 1000 we'd learn a lot. But that's only because we don't have access to much other cultural information from back then, unlike what we'll have now with webpages and news articles.
Excluding an entire medium from preservation would be horrible for future research.
I mean I'm sure it'll be helpful to a subset of researchers. But I don't see why it's Twitter's fault for not making it easy to preserve, as much as it is Microsoft's fault for not preserving WordArt bake sale announcements.
Track the evolution of online language as a consequence of Twitter between the years 2001 and 2100.
Which tweets are just noise now?
I really don't think this is a big problem, in the historical context. Sure, the average person won't be able to casually open those files, but neither can the average native English speaker casually read the original Beowulf!
If a catastrophe wipes out an ancient civilization and all that survives is one laptop-sized cuneiform tablet, historians don't get much from that. If a catastrophe wipes out modern civilization and all that survives is one laptop, historians get gigabytes, maybe terabytes, from that.
Isn't there far more complexity in our current stack than beowulf?
Even if data survived, it would take an incredible effort to use.
I would say yes; we go to great lengths to extract information from hard-to-read ancient sources, so future civilizations probably would too: http://www.bbc.com/news/world-europe-31087746
Compare to the number of people and years of effort that went into building our modern computing platforms. How much research and investment in a very lucrative field? A quick assessment of my university's finances shows they spend twice as much money per diploma on the sciences vs the arts. (Honestly surprised it's that much, though it does exclude medicine which is twice as much as science)
Then consider the hardware and engineers. Just getting to a bit stream is very complex.
I estimate decoding a modern piece of data would be similar to building our entire computerized technology field very nearly from scratch. In my mind that's several orders of magnitude of effort off.
There's also some historical value in journals of every day folks, e.g. soldier's notes from major battles or those present at historical events.
Either way, we're surely capable of leaving a better legacy than what we currently have.
Really simplified... For a letter to be preserved: put it in a box and take it with you when you move and hope that the house doesn't burn down. For a file to be preserved: keep backups to avoid data corruption and media deterioration. Also: control physical media obsolescence and format obsolescence. Most people don't have clue about these things. If our cultural heritage institutions don't get their hands on some of these files, most of them will likely be lost.
And no, it's not only about our great grandparents. It's also about the Turings and Einstens. Their day to day mail correspondence has taught us quite a lot. Turing might not have been a good example though... :)
Paper archives have the unfortunate characteristic of concentrating treasures together. 9/11 apparently destroyed quite an archive of photos - we don't even know what all was there. The Vatican Library concentrates a huge collection of paper, and one little event could take it all away.
With digital, copies are cheap. I've been copying forward my old stuff for 40 years now (though I did lose my old IBM punchcard decks and paper tapes!), and it does get easier. With cheap terabyte drives, one drive will hold it all, and I can make copies to make it resistant to catastrophe.
The Vatican Library really needs to make it a priority to scan all those papers.
For example, I have digitized probably over 5k photographs and several dozen home movies (that I had to record from VHS tapes and even some old film reels), and several thousand pages of documents - report cards, letters, tax forms, etc - of my extend families over the years. That entire archive is now around 20GB big (I have "new" stuff and "old" stuff, pre and post 2000 separate) and I have it backed up in the cloud, on three personal computers, my own external hard drives, and I have it burned in sets of dvds in three places (two homes and one storage unit). I made sure it was all formatted with open standards (extenion-less pdfs, pngs, jpgs, theora, vorbis, odfs, etc). The archives themselves are all UDF formatted.
This kind of stuff is not going to magically become unusable. The formats aren't going anyway. The transport interfaces might, but someone will be able to read dvds, plug usb, and plug sata for several decades at least. And I fully intend to update that collection (probably around 2020) and include the next decade in it in whatever formats are most appropriately modern at the time. At least as long as I still have something like gstreamer to pipe all the audio / video through to transcode it from speex / vorbis to opus or whatever else turns up.
Most families should do something like that. I doubt many of them are.
It rained the day I visited the place, and since it was an old building, it actually rained a bit inside. The guy took it with humour, but it was actually a pity. It is a great place and on some days the visitors can even run old programs on those ancient machines or play games.
From their site: Opened on 22nd of September 2007, Club PEEK&POKE is one of the few permanent displays of vintage computing technology in Europe. Located in the centre of the city of Rijeka and spread across 300m2 of space, it contains more than 1000 exhibits of the world and local computer history, ranging from very early calculators and game consoles to rare and obsolete computers from the nineties.
- you may touch everything, often including things inside the computers itself
- the people working there (or, as is often the case, the one guy working there) is knowledgeable, willing to talk about it in depth, cares, and you can have a real conversation with them.
- there's a lot of surprising hardware around I had no idea existed (from failed computer companies, to specialized I/O of decades past, to special-purpose hardware)
You cannot get the same full experience with most regular museums, as then you have to follow their narrative instead of following your curiosity.
On the other hand, these museums are often not much more than an erudite collection of historical artifacts they got their hands on. To preserve is to select with a plan, not just collect everything.
If you're ever offered to visit the archives/warehouse of a museum, accept. Showing (large) stuff is expensive in space, so most exhibitions are only the tip of the ice berg of the enormous amount of stuff the museum has collected and stored. Often an exhibition is a mix of well-known artifacts that visitors expect to be there (and are therefore included, but are not very interesting as you already know them), some artifacts that are good specimens over-all to show (but not necessarily that interesting), and some truly interesting pieces (you probably never have seen before). I found that when you visit the archive/warehouse of a museum, curators tend to show artifacts from that last category in particular.
Though, the computer history department is a bit too small and similar in size as the equivalent Science Museum in London. In both you find a Cray super computer, first Zuse computers and many older mainframes and terminals, etc. But everything is death, every historical computer sits just there. There is no interactivity, what a shame. They should at least re-work a Cray 1 or 2 super computer and let visitors play around on its terminal - that would be awesome.
My wife and I originally planned to organize a workshop or attend some meetups while there, but the tech scene is so tiny it just wasn't feasible.
We'll be going again this spring, this will be on our list of places to visit.
Maybe it's just that I'm an academic, but this sounds less like a job for a company and more like a job for a library. (Or more robustly, a job for the world's interconnected network of libraries.) The whole point that I take from this article is "we must preserve our heritage for the common good", and that sounds awfully close to a library's core mission.
He's now basically saying "You're all going to lose your data...better give it to Google to save it for you!".
When the guy who wrote the software which produced the file can't view it 20 years later you know you've got a problem.
You can read more about that here:
http://blog.zamzar.com/2012/04/17/open-old-powerpoint-presen...
We shouldn't have to rely on reverse engineering to preserve what's ours.
Unfortunately public administrations don't realize (or don't want to realize) this, at least where I live. I even discussed this with information professionals and they wouldn't understand! We're right in the middle of some big change and it will be too late when they finally realize.
I think the lack of Vint Cerf stories has more to do with the temporal nature of news than any biases that people might have. I'd chalk it up to him just not making news lately. Any text on the history of the Internet, TCP/IP, and networks in general liberally mentions him.
That said, I wasn't aware he was with Google either. I imagine his place in my mind was solidified back in college during the .com days when history internet was a common topic of discussion.
I think that HTML files in a standard character set like UTF-8 could be readable a thousand years from now if human civilization has not destroyed itself.
I hold out less hope for formats like various ogg formats, TIFF, JPEG, MPEG, etc. Software like computer games is even more problematic.
I am hopeful that the technology will improve for archiving digital assets. New storage technologies will become more reliable, much more information dense, and less expensive both to build and provide power for.
Anything involving wireless is a good start. There's a hard physical limit to the amount of data you can cram over 4G or 802.11 spectrum. It's physically impossible to losslessly stream video at 4k 60fps over these, so that's why we use lossy compression (currently MPEG) and always will.
The key to lossy compression is that you can have huge file size savings while not causing visible degradation (visible here being the human eye).
As some of the other posters here mentioned, with the advent of cloud based services and easy backup systems, retrieval is getting easier. So if storage and retrieval is solved, that leaves format evolution problems. But somehow I doubt in 20 years or even more that JPEG is somehow going to be harder to read than it is now.
Other recent (non-photo specific) examples of vast amount of data disappearing:
Geocities. Yes, large chunks of it was archived last minute. MegaUpload.
And we have Rapidshare on it's way to disappearing.
Cloud services can and will shut down, and it is not at all a given that we manage to preserve the data.
This creates a relentless churn where some proportion of older data disappears every day. All we really can do is to fight to keep the churn rate low enough, because we have no realistic prospect of saving everything all the time.
"The project was stored on adapted laserdiscs in the LaserVision Read Only Memory (LV-ROM) format, which contained not only analogue video and still pictures, but also digital data, with 300 MB of storage space on each side of the disc. Data and images were selected and collated by the BBC Domesday project based in Bilton House in West Ealing. Pre-mastering of data was carried out on a VAX-11/750 mini-computer, assisted by a network of BBC micros. The discs were mastered, produced, and tested by the Philips Laservision factory in Blackburn, England. Viewing the discs required an Acorn BBC Master expanded with a SCSI controller and an additional coprocessor controlled a Philips VP415 "Domesday Player", a specially produced laserdisc player. The user interface consisted of the BBC Master's keyboard and a trackball (known at the time as a trackerball). The software for the project was written in BCPL (a precursor to C), to make cross platform porting easier, although BCPL never attained the popularity that its early promise suggested it might."
http://www.bbc.co.uk/history/domesday
And I believe the content on the original Laserdiscs has been reverse engineered more than once. Hopefully once content is on the web (unless it's behind robots.txt) it gets sucked up by the Internet Archive.
That of course doesn't take care of any rendering issues.
Really? Most of the big supermarkets here (UK) and quite a lot of pharmacy chains have machines where you connect your device/USB/SD, select the pictures you want, and press print. It's relatively inexpensive too and I know people who use them all the time. Personally I don't actually mind if all of my digital photos are unreadable so long as I have the important ones in a physical format so this is a nice solution and takes away the pain of owning a good printer, buying lots of ink and photo paper.
https://www.google.com/settings/u/0/account/inactive
In the event your google account becomes inactive for a long period of time you choose, you can trust specific contacts with access to specific data. It has other niceties. I highly recommend setting it up.
The question is what will happen to Picasa? Will it be free in 5, 10, 50 years? Will it be free?
We still had jpegs 10 years ago. All of my digital pics from 2005/6/7 till now are still on multiple hard drives and multiple systems. Each time I get a new computer I transfer them over.
Why are they unreadable? It sounds like you put them on a proprietary format or some obsolete picture software you used to burn them on.
I think its absurd that you are paying for picture storage and think its your only option. Amazon prime customers can upload pics for free now [1]. Why not use one of the free ones like Dropbox, Google Drive, SkyDrive etc...
CD-R's, and especially older CD-R's, turned out to not have as long a shelf life as initially assumed. I recently found a bunch of old CD's I burnt ages ago, back when I got my first CD burner, and only one of them was still completely readable.
My parents have CDs from the 90s that still work.
Would there be a difference in a music CD you bought from a store and one that burned yourself?
There are also special archival disks and burners, like M-Disc, which literally "engraves in stone": http://www.mdisc.com/
All of them have space limitations - and I have over 400GB of pictures backed up to Picassa. So if I am not paying google, I would be paying Amazon or Microsoft or Dropbox or someone else. And my main point was - what happens to data on all of these paid-for services once I am dead? Amazon Prime backup is "free" as long as you keep paying for prime.
And yeah, CDs become unreadable after sitting in their envelopes for 10 years. Not all of them,but mine certainly did.
The status quo is probably stable for games and big name software packages for the time being, but it doesn't inspire confidence for the long term.
It's always been hard to migrate out of social systems--convincing a well-connected user of a photo service to move away from the place where they've accumulated comments and tags is really hard. That metadata is not portable because identity is not yet portable, and it's what we're spending time on (lots more time than we spend making spreadsheets).
I think we might manage to keep "JPG as file" alive for 30-50 years, but there is lots more to manage.
We are getting better at archiving files. And thanks to a combination of people dedicated to emulation and the rise of virtual machines, there are few popular pieces of hardware we can't emulate in excruciating detail.
But many of the services I used a decade ago are already gone. And pretty much all the services I used two decades ago have completely disappeared, with some very few notable examples. And with them, vast amounts of data.
Some of it the Internet Archive have at least captured static snapshots of (and they really should have magnitudes more funding), but ten times that - or more - was data in walled gardens, behind logins or otherwise restricted in ways that means it is lost forever unless we're lucky and it turns out some admin held onto backup tapes they weren't really meant to keep.
And the problem with there is not to create a snapshot of a single server, but as you say that distributed systems are far harder. Even recent. Twice I've been contracted to help companies take over infrastructure that involved systems I'd worked on, and try to "package it up", and it was incredibly hard, because no matter how much you try to tear down and bring up individual servers or groups of servers and automate deployment, very few places running complex services ever try - or could afford to try - to tear down and bring up a full copy of their entire infrastructure.
Suddenly all kinds of nasty interdependencies and bootstrap problems nobody had needed to think about shows up.
Text documents in obscure formats can be more troublesome. There were many early word processors, and many file format versions. Those can be hard to convert, and there will be obscure text documents some historian will want to see a century from now. Converting stored text documents into some self-explanatory form like XML for archiving purposes is helpful. Even if the software doesn't survive, the text will still be there and someone can probably figure out the encoding.
Structured graphics files from old programs are a real problem. This is a big problem in the CAD world. CAD files aren't just pictures any more. They're detailed descriptions of physical objects and how they're made. Moving them from one present-day CAD program to another is tough. Going back 20 or 50 years will be tougher. People will have a real need to do that; buildings, aircraft, and industrial machinery last that long. The present compromise is to export such things in well-know formats that are viewable, but not necessarily editable.
Cerf is talking about preserving execution environments, so you can run old software years later. With so much "cloud based" stuff, and network oriented DRM, that's not going to work once the servers have gone away.
Imagine you're an archaeologist in the year 4500 or so, by our calendar. What we know of as modern Western civilization collapsed thousands of years prior, all you have are physical artifacts dug out of the ground in your attempt to reconstruct the history of this lost civilization. What would you see?
Circa the late 20th century you'll notice a precipitous decline in the volume of any surviving cultural material--meaning printed books, magazines, business papers of various sorts. You'll find fewer sound recordings, ticket stubs, even purchase receipts. The various detritus of daily life will appear to rapidly dry up and virtually disappear, around the world more or less simultaneously.
What conclusion would you make? The population at the time seemed to be stable or growing as measured by ruins of settlements, yet they seemed to be doing less or producing less? Perhaps there was a crisis in education and illiteracy became rampant? Maybe there was an ecological disaster and paper itself became scarce?
You see that a huge proportion of our culture that we take for granted--not just pop culture, cat gifs, etc--but substantial business and scientific research information as well would be completely lost to these future historians. On a more trivial level, when was the last time you saw a comprehensive, printed guide to iOS 8 (for example)? You could consider that a significant cultural/artistic artifact from our time, yet how will it be preserved in any meaningful way for future scholars?
Anybody interested in these questions might peruse http://www.longnow.org
1) A future where society loses the ability to build and maintain tools to process digital data at scale, and as such, are only reliant on analog tools to reconstruct the past.
2) A future where due to the extreme advancement of technology, old file formats are "forgotten", and therefore they are difficult to decode.
I think Vint statements are to be taken in the context of scenario 2, although I kind of think that even if we forgot the format, we'd be able to rediscover it with enough analysis and computing power, which would be a non-issue in future scenario 2.
For scenario 1, I don't think we have answers short of burying long lasting dormant computing devices in bunkers all over the world, along with maps of how to find them cast in a very stable medium.
Or see http://rosettaproject.org/disk/concept as a proposed analog preservation solution for a large(ish) data set.
Sure but at what rate? Most 20 year old CDs are still fine, to say nothing of paper stored with a sliver of care, what percentage of web companies from 1995 are still around?
This also leads to the conclusion that a standard mechanism for interfacing to the data from multiple apps on multiple data backends is needed. The Android storage framework is probably the best effort at this so far, but it's far from clear how used it is.
But seriously, these complaints are dubious at best. Changing hardware? Well, maybe. Still, we won't migrate to newer hardware unless we can bring our stuff with us. Maybe only if 2D pictures will eventually be considered obsolete and never used since, but I doubt it as well. Changing software? I'm struggling to imagine how text documents could become unreadable. Even on completely new architecture it won't be hard to write a translator. The same way I cannot imagine bitmap images becoming obsolete, and every single format we use is just moderately complicated compression algorithm wrapped around bitmap, and every curious historian will be able to recreate it by himself. The same stays true for wav/flac,ogg,mp3, etc.
I can imagine how Adobe swf will become obsolete and it might be hard to find software to open it in 100 years. Or Microsoft Office slideshows. But it feel almost right.
Encrypted stuff maybe totally unreachable though besides brute force.
Which is why I think that the resuscitation of 8-bit Computing, specifically, is such a valuable thing to do: it provides context. When you've spent the evening actually having fun with 30-year old software, the urge to splurge on soon-to-be-redundant newgear is de-composed. Eventually, a person can understand that all computing architectures over the Age, So Far, are of use. That's how they got to be a working program in the first place: someone found it useful.
I recently downloaded a PDF of 80 or so BASIC programs, written to be as compatible with the plethora of machines that were available in the 80's, as possible. What a joy it was to see linked lists, self-modifying code, and competent optimization of program space while also using simplified interfaces, to be cross platform as possible. A modern comp-sci student can even still today, learn a lot of very important lessons about computers by reading such archives and going through 30 or so years of history. It factually is not a long amount of time. All those lost floppy disk collections, out there in the dumping grounds, or even the ones still working, hidden in the closet, have the potential to be just as relevant in 100 years as they were on the very first day of publication.
I urge anyone with an 8-bit stash to dig it out, soon enough, and find your active community. There are few 8-bit machines out there which don't have a thriving scene.
By essentially saying that, if you don't use cloud technology your precious photos and documents might be unaccessible in years to come.
http://www.everpix.com and the like. Hopefully less often than hard drive failures, but more frequently than a format like JPEG becoming unreadable.
I really hope someone comes up with a good long-term data archive/retrieval system soon, but I'm not holding my breath.
"Never trust a corporation to do a library's job" https://medium.com/message/never-trust-a-corporation-to-do-a...
An example of acquishutdown from HN: https://news.ycombinator.com/item?id=8472047
I have no idea what the rate of decay for digital records, information, and archives would be, but I would think it would be higher than information stored as hard copies of paper, books, etc. Of course, we are also producing orders of magnitudes more information than we were in the past.
I would argue that for the first time in history, we're not limited by storage space.
"We know why the dark age happened [...] Our ancestors allowed their storage and processing architectures to proliferate uncontrollably, and they tended to throw away old technologies instead of virtualizing them. For reasons of commercial advantage, some of their largest entities deliberately created incompatible information formats and locked up huge quantities of useful materials in them, so that when new architectures replaced old, the data became inaccessible."
This puts another perspective on the story about Japan's oldest companies, does it not? (https://news.ycombinator.com/item?id=9041040)
Why not regular servers?
Firstly, the medium of storage, the encoding and the interface form one angle to look at. Just like how floppies and CDs are now (almost) obsolete, there would be a future when there won't be any machines that recognize a USB storage device. The same would hold good for other technologies that we have quickly run through, to mention a few - PATA, SATA, PCI, PCI-X, PCIe and so on. Can you read an MFM hard drive today with any computer that you have? With adequate care, data can be moved from one medium to another, like how people learned to move music from tapes to CDs and then to flash drives and hard drives, then to the truly nebulous thing called "cloud", etc.
Next, consider the data formats themselves and the applications that support them. This is where proprietary formats, especially those that are not widely popular, would hurt the users. So if you're using, for example, Apple's document formats on Pages/Numbers/KeyNote, it's likely that those files will soon become obsolete (as they already have, where Apple does not support older formats in the newer iWork). Commonly used and supported formats that don't change rapidly, like JPG and PDF, for example, are safe for a much longer time because there are many applications to process these documents with, both proprietary as well as FOSS. Even the web pages stored by archive.org or any doc or xls files lying around from about 20 years ago - how many more years do you think browsers, word processors and spreadsheet programs will keep bloating up (like they have been so far) just to support older versions of doc, xls, ppt, html, etc.? At some point, the bloat will have to be cut and a decision made that older formats will not be rendered like they used to be. That would leave some clobbered text or gibberish or both showing up for any interested future humans to figure out whether it's worth preserving or not and to convert it to a newer format if they care.
Now, assume the data format is a long surviving one, like say, mp3 or jpg. How do you protect it from bit rot wherever it's stored? Just because you put it on Amazon or iCloud or Dropbox does not mean you can't lose part of the data to corruption of different kinds or lose all of your data due to system failures. Among the people who do regularly backup, perhaps only a fraction of a percentage actually verifies backups (if at all). With consumer level cloud options, there's not a lot of hope for data longevity for non-tech savvy people.
If you're dreaming up a beautiful future in the cloud for all data storage, how can you be sure that your data just doesn't vanish or that it doesn't diminish in quality? People have had that happen to their precious photos by sites that had sneaky terms and conditions about deleting old photos (or if photo prints are not ordered regularly), didn't allow full downloads at any point in time, and sites that went out of business. We can see people treating social networks as reliable cloud storage for their photos as well, disregarding the risks of making such an assumption.
Ignore data integrity and data format issues for a moment, and imagine a distant future where 64K displays are common. All your current and older photos and videos would either look terrible on those or may even be completely indistinguishable. Of what use would these artifacts be then for anyone?
Considering that a lot of data on the cloud is actually insignificant at the level of the human species (like rants and comments on social networks, LOLcats, etc.) and also the huge amount of data being created every second, how would someone in the future even sift through all this? It's somewhat similar to the NSA/GHCQ looking for a needle in a haystack the size of a huge mountain, except that this haystack is to preserve history, culture, etc. The current haystack being built would need a lot of archivists from different backgrounds working continuously to separate the wheat from the chaff and to also look at consolidating (or "packing") the information concisely (like gathering summaries, sentiments, trends, statistics) if we're ever to have an archive that future generations would even want to look at (say centuries ahead in the future). Leaving it to governments alone or corporations alone is not the solution since each would shape the archive in its own image.
This is a very complex topic for most individuals to deal with, and the above points didn't even touch upon cultural and linguistic shifts that happen over time for any data to be usable. I'm sure I've missed many other aspects about prolonging the life of (usable and useful) data.
P.S.: The best everlasting format, as many tech savvy people know, is plain, unencrypted text for textual content (this still assumes that the media can be read, because encoding may play spoilsport).
P.P.S.: All LOLcats may actually be cultural items to preserve for eternity! :P
Now get off my lawn!