Shining a light on the digital dark age
longnow.org
longnow.org
ETA: The issue I face is two-fold. More pressingly, there are files just missing. Nothing I can do about that barring someone magically having them downloaded to an old PC. But also interesting is working with old formats. The 2000s/early 2010s Halloween Horror Nights websites were all written in Flash. They have tons of little Easter eggs and information for the event obsessives like me. Between fan backups and the Internet archive, the files are pretty complete thankfully. But since Flash has been dead for a while, I have to rely on the Ruffles Flash emulator to get it running on the web. But that doesn't work super well. On my list of things to do is to contribute to the project to try to get some of the files working.
In case anyone is curious about my backup of the event or if you're also a fan and have any of the files listed to share and archive, my site is hhncrypt.com!
Even a single-file website with 2 photos can be important to somebody. Thanks.
The new music is fine, everything is on youtube.... for now! ... but the older songs are disappearing.
Years ago, we had a bunch of "mp3 sites" (websites where you could download pirated mp3s), that disappeared, a huge local torrent tracker, focusing also on local stuff, that lost most of the data in the OVH datacenter fire and slowly died, and some stuff was put on youtube 15 years ago, and then got copyright claimed 5 years ago, and was removed. The music groups don't exist anymore, so the CDs aren't publushed anymore, second hand shops are very rare and cater mostly to LPs and tourists, and the groups stopped existing before they could sign deals with the publisher for streaming services.
So yeah, it's not just "those semi-personal photos taken on a party and kept by maybe one or two people", but also pop-music that used to be on the radio all the time in late 80s, early 90s, and is just..gone!
I'm sure that there are data hoarders somewhere, that have mp3s of all those songs somewhere, but unless we get some p2p type of service like gnutella/ed2k/kazaa working again, I won't be able to find them anymore.
Such archives would need to be kept online though. If only archived (or hoarded), it's just data in a box that no-one can access.
Current internet is really poor at handling this long tail. Eg. a movie can be streamed & torrented, millions see it, many have it in personal archives, but 1 year later the torrent swarm has died out, and those personal archives of the downloaders aren't online. And then copyright holder pulls it from their streaming service.
Result: nowhere to be found. Even though popular not long ago.
Copies in personal archives don't count for much if others can't access that data.
Eventually, after searching on and off for a few years, I found a copy on a Russian streaming site, but the movie has effectively been erased from availability.
My philosophy around this is "File over app" — if you want to create digital artifacts that last, they must be files you can control, in formats that are easy to retrieve and read. Use tools that give you this freedom.
In the fullness of time, the files you create are more important than the tools you use to create them. Apps are ephemeral, but your files have a chance to last.
I now use iA Writer which allows you to save your notes in markdown from iPhone, syncs via iCloud, and then I can continue from my laptop. Admittedly the native Notes app has some better functionality around sharing and searching. However, Notes saves into a SQLite db, which would be fine, but it’s not trivial to view the tables/schema in there.
Slack’s limited search history is a feature too - forces you to document using appropriate tools and not endless email threads…
I'm sure that I don't know what you mean. My employer is on a paid plan for Slack, and searches cover everything, as far back as I wish to go. Are you thinking of the limitations on the free license?
The other great thing about slack is that anyone from company can start an account without IT’s approval…
This reminds me of books. I’m sure the majority of books from over a hundred years ago are lost because they weren’t popular. We haven’t really noticed their absence…
Especially if you include independently published books that weren't widely circulated. I wonder what percentage of total books this is.
My grandfather published a book before he passed away. It was never sold online or in any big retail stores. Once the last hard copy is lost, it's gone forever.
i believe the Library of Congress will archive that book for you if you mail them a hard copy. assuming it has a ISBN, you’re in the US, etc.
Not all books in the state libraries are equal. Historical copies and popular authors (popular among researchers, a much bigger set already) are exhibited and get attention, John Doe's book of family recipes gets sent to some giant dark warehouse people rarely visit.
It is easy to forget that it is an 18th century solution born from 18th century approach to knowledge. Back then, bibliographies of everything printed in certain year in certain country could be compiled, and they were supposed to be more than just lists, to help other men of books keep up with Progress.
So, 10 hours of white noise yes, some person's personal blog where they poured their heart out, no.
Beautiful.
Kierkegaard, Thoreau, Dickenson and Melville, for instance.
If their works had been lost, "we" probably wouldn't have noticed any of those absences either.
If they were in one of the university libraries that Google scanned, they're not "lost." But you're right; you can't read them. Congress should mandate that the Library of Congress, at least, get a copy to preserve them for the ages.
Read the Atlantic article
https://www.theatlantic.com/technology/archive/2017/04/the-t...
for the sad story.
For example, if LLM NN model weights are distributed with IPFS instead of corporate infrastructure (basically zero redundancy) the popular models would be very available, and have essentially near zero chance of being lost.
To state that again, the llama models likely have tens of thousands of downloads, which would mean tens of thousands of servers and backups of the data, versus what we have now, which is essentially just one.
We need IPFS for data distribution. Tightly knit integration with git repos is an obvious match as well.
Using today's technology, it is possible to allow exporting all content as basic file formats such as txt + zip.
In addition, PKI (public/private key infrastructure) allows us to decouple a user's private and public identities, meaning the public identity can now be portable between servers.
What does all this mean for the average user or community operator?
It means your community can be completely transparent, auditable, and PORTABLE, allowing any user to archive the whole thing and clone it to another server -- right away or years later.
I've been writing a framework for this type of system for several years now, and if you're curious, you know where to look.
Thank you for coming to my ted talk.
I think part of the problem is that most backup efforts result in lawsuits. Related, I wonder how much the data stores, like those that OpenAI used, have preserved, and I wonder how much they will purge from their servers, to be lost forever, as the lawsuits increase.
For Youtube, it also has compression rot. It's not a storage solution, as I've learned. All the videos I uploaded have slowly reduced in resolution, bitrate, and quality, over the years. Those that are more than a decade old have become a blurry mess that I can barely see, at a fraction of the resolution. I can't blame them. They don't really have views, so they're a money sink.
i know that feeling when i watch "old" stuff too, did you confirm this in any way or is just subjective?
asking because after a while of watching pal/ntsc style content it becomes normal again.
.
Generational churn means eventually no one who knows why the nested dolls were nested as they are will be gone. Humans will create a new set of nested dolls they can grok.
Paraphrasing Thomas Jefferson; clearly the dead do not rule the living.
Edit; forgot this point… Maybe he said that; it’s unverifiable for us. Maintenance of hallucination is all it ends up.
Reality we see, smell, hear, and touch is what we get. There’s no violating physics. Let it be lost. It’s going to happen anyway.
Yeah but don't listen to him, he's dead.
Hence we can safely conclude that the dead do rule the living.
Checkmate, Epimenides!
Without direct observation it’s hearsay.
What’s the point of maintaining associations we can’t verify? The truth that Jefferson stated such is only hallucination for us.
The value is in the awareness of the realities mechanisms, not the association with Jefferson.
Obituaries that appeared in print newspapers during the 20th century were easily disseminated and decentrally archived (typically by loved ones and libraries), making them relatively rot-tolerant.
Distribution isn’t a problem for digital obituaries, and in many ways the web is better than print in this respect.
But when it comes to preservation, there are many factors that make digital obits in their current state particularly susceptible to rot. They tend to be centrally archived and often behind paywalls, making them susceptible to digital rot and difficult for organizations acting in the public interest to archive.
The for-profit company Legacy.com controls a strikingly large share of the market for digital obituaries. It partners with funeral homes and newspapers, and in many cases when a visitor browses obituaries on the website of a local newspaper or funeral home, they’re actually redirected to Legacy.com, which hosts the content.
Unfortunately, the newspapers and funeral homes themselves often don’t maintain their own copies of the obituaries. What happens if Legacy.com or one of the smaller memorial sites goes out of business or experiences some sort of data loss? Because of the centralized nature of how these digital obituaries are stored, it’s possible that very few other organizations will have archived copies of the content.
Consider that everything you witness has an expiration date and if you don't save it maybe nobody else will.
Part of storing data is the cost per byte per year. That cost is ridiculously low on my home 52 TB workstation. (2 TB NVME, 14 TB and 2x 18 TB HDDs)
A backup copy is stored on 3x 18 TB HDDs in USB enclosures.
Example: my Netflix data exports from two years ago include much more detail than the recent ones. Netflix is just now providing overall less details.
It's why I don't delete some of my accounts when I think I might be able to get more data in the future from some company
The government is probably already doing with the NSA but we normal people can’t access it.
Even in categories such as fiction, there is no reason to keep every novel written. Only a tiny percentage have enduring literary value.
For society, as for the individual, forgetting is an important process in maintaining sanity.
I can rightly say I am a Diplomacy tournament champion.
here's my problem. The 1st, that I won, was during The Digital Dark Age. around 1990
The 2nd, the no-prep walk-on that I lost, was like 15 years later. When everything was getting put on the web somehow, Google-able and archived etc
my Diplomacy record therefore is Heisenbergian. I am a champion... except if you Google to confirm it. in which case you'll see that I lost
both things are true. paradox FUN caused by the Digital Dark Age
If it happened before the internet, or during an early stage of it, or there isn't an archive of it somewhere online, did it really happen??
(Cough, someone still has to store the information.)
Oh please. You'd have to be naive to think otherwise.
Specifically, I've been recently looking for accounting software for Polish VAT accounting.
I'd much prefer standalone software rather than SaaS, but I found none. Only SaaS offerings. Of a dozen of them only one has a public API one can use to export data(if one writes exporting software). None have ability to export all data in a simple format in one go.
Why does it matter? Accounting records, buy/sell invoices, expense documents etc gathered over years and decades can be very valuable.
Personally, despite being against over regulation in general, I think there should be some law that required SaaS companies to provide reasonable data export capability. GDPR already gives people an ability to request all data another entity has on them, but making it an essential part of the service that could be used regularly (as a backup of sorts) would be much better.
I made sure the backed-up copy on my machine was ok. Then I purged my bucket and am proceeding to delete my long-stagnant account.
Can't fault them for trying to make money but it sure seems capitalism is a big fan of entropy. And it's good to remember that you see your data as an asset, while it's aways a liability for a company providing services; free, paid, or otherwise.
If I didn't happen to have the same email, I wouldn't have received the account deletion email to even check what was in there (and reset my password). A whole lot of early social internet content is gonna get poofed quite soon.
I'd say we should sign everything we can now, anything created from 2023 onward is already suspect of being created by AI. The past will be 'erased' as well if there's no way to verify our historical information.
For example, I create an AI photo of Frank Sinatra in an LA diner eating a sandwich and post it online - tell me how on can verify today that picture is legit or not. Whose the arbiter of all Frank Sinatra photos? How much time, effort, and money would it take to do that verification? Now extrapolate this example to everything. The past becomes only myth and legend.
Also they usually like to burn any -history- or -culture- of the loosing side and adapt their customs to their new ones and call it a day, erasing history pretty much, as much of it as they can at least.
YMMV
- You usually can attribute history written by a victor to said victor;
- There's only so much control a victor has over what's being kept by the monks, librarians, museum curators and individuals, and what of it will resurface once they're gone.
With AI, we're not talking about alternative history, but rather about infinite, arbitrary alternative histories that can't be told apart from the real one.
I can imagine a couple million years from now, some alien species shows up, we’re all gone and they think maybe we had wings, some of us were born with blue hair and other were half robots. I get they can study some of our remains, but so much of us is mutable digital info now.
Historically that has not been the case for most of human history. As odd as it might seem, in general the rediscovery of the accomplishments ancient world has been a great driver towards progress. The periods when the accomplishments of the past were lost and fully forgotten were the sorts of times people call Dark Ages.
Obviously we can’t be sure how the current era will be viewed from the far future, but your comment made me realize that the current situation has similarities to that dark age.
(or the ones where the "old city" was destroyed in brutal urban warfare)