I remember a few years ago TI got bored with updating their cpuwiki, and deleted the whole this - which stings to this day. The internet is still full of links to it to solve crucial issues, only to be greeted with a disappointing error message.
I remember a few years ago TI got bored with updating their cpuwiki, and deleted the whole this - which stings to this day. The internet is still full of links to it to solve crucial issues, only to be greeted with a disappointing error message.
As a sibling comment says, I think such sentences have meant "you can't rely on the internet to forget". Of course the reverse is also true: "you can't rely on the internet to remember".
In short: don't write things you could regret, but make backups.
The internet does forget in the sense that stuff goes missing. Just, only stuff that you don’t want to go missing!
It amazes me how companies will have free volunteers help people to use their (often expensive) paid subscription products, and then delete all that info those volunteers wrote up. Don't they want people to use their products?! They're less likely to renew their subscription if they struggle or are unable to use the product for their particular use case.
Unaffiliated forums not ran by the company are better in that the company can't decide to just delete all old posts one day (and while the owner could, certain types of unaffiliated forums are usually a bit easier to clone and republish.) The downside is you don't get assistance from people who work for that company, but often you rarely get that in official forums. The usual reason to use official forums is just that they have significantly more users asking and answering questions than unofficial ones.
I dream of someone taking the internet archive data, capping it at 2010 or so, then making a search engine out of it. I mean if AI companies are looking to gobble all the data they can get, then surely they'd jump at the chance to train on (higher quality) data from the past that simply no longer exists on the web. So it'd seem like a win-win situation if IA gave them a copy of the data on the condition that they maintain a permanent backup and provide some sort of searchable index on the data (maybe even via LLM), and in turn the AI companies got access to high quality data on obscure topics that simply no longer exists.
It sounds like you're describing CommonCrawl.org, and yes, it's already popular with AI companies.
Personally I think it is a huge error we don't force archive.org to allow others to mirror their data easily.
What do you mean by this? I thought they definitely were open to mirrors, why would ’we’ need to ‘force’ them to anything?
IIRC the Bibliotheca Alexandrina in Egypt had a mirror of the web archive up, though they may have failed to maintain it
I can already look at e.g. the National Film Registry and say I don't care about movies so my tax dollars shouldn't go towards preserving them, regardless of what some bureaucrat thinks is "culturally significant." The film industry has a pretty massive network of artists who are able to go to bat for that stuff, though.
But then take pornography, or hate speech (and/or straight-up misinformation), or content that's illegal in the US but not elsewhere (like drug recipes or weapons schematics). It's already tricky for a private entity to handle convincing people of the value of preserving everything physically and legally possible to preserve. Making an argument for doing it as a public function, even in good faith, could easily turn into a mess-- especially when we can barely agree what should be allowed on the internet now.
If anything, governments should be proactively funding organizations that archive content and governments can archive content of significant cultural or historical value like the library of congress does for physical media.
I mean we wouldn’t want companies to fall afoul of some law because they had to take their forums down due to some privacy ruining bug in the software. Or because the old forum server sitting in the basement died. Or because the third party software they used for the server went out of business.
With the exception of maybe a transient bug, isn’t this exactly the point?
“We wouldn’t want a company to run afoul of [data preservation law] because they [neglected maintenance]” seems like the directly incorrect intent. We would want to compel the business to migrate or modernize any hosting to keep it active and viable. If their vendor goes out of business, they should’ve paid more or migrated the data.
If we’re discussing a local club then sure, they’re a victim to hardware failure or business changes, but this thread is full of billion-dollar-businesses. They can spend a few thousand dollars on a forum every few years. When I worked at $FAANG, my service had millions of users and cost like $10K/mo in hosting. Surely the Autodesk forum in read-only mode would cost a much less, and almost nothing if migrated to static HTML.
There are companies that have a forum, but it just seems to be a best-effort mostly community driven thing, or at least it isn’t tied to any paid product. It would be a shame, IMO, if they couldn’t offer that without signing up for some perpetual obligation. Even if it is small, somebody has to have an eye on it.
The web archive already works from the mode that anything can disappear at any time. Anyone who would rush to archive it if it were going offline could have archived it any time over the last N years.
Websites getting shut down no matter the reason is just a part of life that we need to accept. Laws can't address the underlying churn of time. Forums should be easy come easy go, not open you up to litigation if you shut down your own site too quickly. Cmon.