38% of webpages that existed in 2013 are no longer accessible a decade later
pewresearch.org
pewresearch.org
If a business is only on Facebook, I don't do business with them as I don't use Facebook.
A win-win in my book, as I prefer doing business with people whose ethics overlap with my own.
But it's clear that continuing to use Facebook in that manner will only strengthen the isolation effect. Voting with your wallet and going against the machine invariably involves some level of personal sacrifice. For me, sacrificing patronage is incredibly easy to do. There is more to life than commercialization.
My girlfriend says she only uses Facebook to interface with small stores, who use it as a sole point of contact or distribution. Let that sink in for a moment. Breaking this cycle will require hard work.
I have avoided a place with only insta, simply because i couldn’t see anything I needed to.
You used to be able to see a custom feed of a selected friend lists but since they removed that option the site has been completely unusable, unless perhaps you do something like remove 90% of your "friends" and groups but that would hurt usability in different ways.
I liked more the Facebook that showed me the humblebrag posts of my friends/connections (and I'm not being sarcastic)
The topic is FB groups. They aren't spam, at least for those I'm a member. Some groups may be quiet, some are active, but I don't recall coming across spam posts from any of them. A particular group has a rule that members can promote their business once a week, enforced by the group's admin
Not only a lot of communities are hidden because of Discord (at least with Reddit they were more discoverable), the worst part is the fact that they are unsearchable or behind a paywall.
Like the "join my discord if you pay at least 3$/mo!" is pretty innocent but you are gatekeeping a community that before was pubblic.
If we are talking about something like a content creator focused about an hobby or pc problems you can see how Google will become even more useless.
Reddit was the least bad choice between it and Discord but has failed the "i want to be a social network".
The big thing about discord is you can chat now with people but the knowledge is not in a good format to come back to later.
However, if you managed to find the forum on a search engine and took the trouble to sign up for an account, you are more likely to abide by the general vibe of the place, rather than Redditize it with shallow, meme-y comments that reliably get a lot of upvotes.
Independent forums (phpBB, and the like) often came up on searches before this communities moved to Facebook Groups, where they’re mostly set to private due to spammers.
Similarly there was a time when Google indexed tweets more or less live, so you could find information for very recent events. I think Twitter asked for money and so that was the end of this.
Now I think Reddit, and maybe Stack Overflow, are the only things helping Google be anything more than an extremely hostile version of the yellow pages. I fear Reddit might at some point withdraw their content from Google and that’ll be the end of it.
I upgraded to VPS for $500. The other admin spent 15-20 hours fixing/troubleshooting/transferring. And you know what? At the end of all this, I paid to give my data to these jerks, to keep it online for them to harvest. The forums are dead quiet.
Now I think, Discord is fine. They'll just sell the data to AI companies directly, the burden won't fall on me.
The open internet seems increasingly predatory and a place where some gigantic ML company just vacuums up your stuff or resells your content for ad revenue, parasitic.
I don't mind the fact and think it's honestly a natural reaction to this that people guard their information. It's sort of like a medieval monastery version of the internet where people recognize that information is cultivated rather than just some commodity you scrape off the web.
I think many of us (wrongly) have a tech-centered view of online communities -- witness the multitude of "Show HN" posts that are "look! I made an online place for humans to congregate and discuss X".
The tech stack matters little (if at all), but bootstrapping the community's trust and culture (and maintaining both) are most of the heavy lifting, and the differentiator for success.
I recently (re)discovered an HN post by one of the core community moderators of deviantArt, and its success was made possible by its culture:
http://news.bbc.co.uk/hi/english/static/in_depth/americas/20...
http://edition.cnn.com/SPECIALS/2001/trade.center/index.html
Don't expect many of the links to work properly, but it's still interesting to see what the web used to look like.
Hard to imagine that with many sites now 20 years on. It's not even that it;s impossible with the technology, it's probably closer to how writing got worse after the invention of the word processor. Every thing is managed and structured now so the freedom / bubble needed to make things good in a way that can't be easily explained is gone.
Bookmarks are good only for links to things for which only the most current version is worth accessing. That’s my banking websites, a shopping site, my employer’s remote desktop system, etc.
I like the idea that in addition to saving the page, you can annotate it as well.
javascript:(function(){daurl='https://web.archive.org/save/';if(location.href.indexOf('http')!=0){input=prompt('URL:','http://');if(input!=null){location.href=daurl+input}}else{location.href=daurl+location.href;}})();
This has the keyword "way" in my Firefox bookmark.So when I'm on a page I want to save, I just press Ctrl-L (to focus the URL bar), then type "way" to save the page on Wayback Machine.
What would be particularly interesting would be to graph how long websites last for. I suspect quite a lot of the content from the early days is still around, and this period (2008 - 2018) is the peak of sites vanishing.
- Geocities
- University-provided FTP folder (deleted after you graduate)
- ISP-provided FTP folder (all those Earthlink, Juno, Comcast sites: probably deleted)
Despite being in 4th grade when my little friend and I made the webpage, things on there (while fine for the era) are just not okay by today's standards even if I understand the context for what led to it being there. It was nothing terrible, but just distasteful in a blissfully unaware way a 4th grader in the 90's would be. I realize that stuff will probably never be off my conscience and I just have to deal with it and hope nobody sees it.
Thankfully, even the archive occasionally takes stuff off.
I think it is bit of a shame that mirroring is not more readily built into web stack (=http/html); if you could trivially make links that included local copy (as fallback?) this linkrot would be far lesser concern. The way how for example wikipedia links everything through archive.org is bit of a hack imho
https://nuim.libguides.com/referencing/DigitalObjectIdentifi...
But it would need to handled automatically to maintain the full usefulness and convenience of URLs. Not sure how that could work though.
This is an orthogonal concern, and arguably is mainly about privacy
> It’s also good that some preservation effort is necessary for worthy content: the value of it gets more appreciated.
This same argument seems to imply that virtually everything should be expensive. Cheap storage is bad because we don't appreciate the value of the files we store. Expensive healthcare is good because it really makes us appreciate our organs.
> worthy content
The hard part is looking into the future to determine which content will be considered worthy then. So far no human civilization has managed to figure that out. They mostly seemed to focus on preserving the image of how amazing their kings were.
Non-zero doesn’t mean “unaffordable” or “expensive”.
> The hard part is looking into the future…
The hardest part is to understand that the content we want to preserve carries more valuable information about us than about itself.
Scientific knowledge can be discovered again, it’s not something to worry about. The preservation shapes the future views of us, leaving the trace in the history of those who preserve, their life and their experiences. Maybe they just needed to accept their mortality and irreversible flow of time?
There is certainly a cost of storing data, and cost should enter the equation. But we're losing a lot of data for reasons other than cost and we don't have a reasonable way of assigning a value to the lost data.
If man was content with the nature of things he would never fly, or go to the moon, or any of the other myriad accomplishments humanity has made. If we can preserve clay tablets from thousands of years ago we can find some way to keep the information we produce today for posterity.
What's the proof for Fermat's Last Theorem again? Doesn't matter, it was just a footnote anyway so let's not bother preserving it. It doesn't matter that it took our smartest minds 358 years of trail and error to rediscover the proof. It can always be discovered again.
Yes, rebuilding civilization from scratch would be a difficult task, taking centuries if not millennia, if no knowledge is preserved. However we do preserve it and do spend considerable effort, what cannot be said about our culture and individual experiences.
Let's keep making preservation easier, and preserving as much as we can. Maybe much of it is worthless, but I guarantee there's at least one document we think is worthless today, that historians 500 years from now will be glad we preserved anyway.
It is highly subjective. I’m very curious about the past, but I don’t care if nobody will know the name of Newton or Mandela in 10000 years, but some YouTube blogger will somehow be a legend.
> morally correct
How can enthropy be morally correct or incorrect?
You said the disappearance of content was a feature, not a bug. If it is a feature it was designed. I understood your comment as implying that somebody created this feature.
You now speak about entropy, which one? Boltzmann's or Shanon? This doesn't have to do anything with bit rot or the like. When you write a book, you cannot unwrite it. Is a fait accompli. But if I create a website and I load a bunch of content and after few years I don't pay domain, server space, etc. it will be deleted; at that point a webcrawler may have had copied all the info, or not, we do not know, or somebody may have thought it was worth it and saved a link to it. At a fundamental level, who decides what stays and what is deleted, who owns it if the webcrawler stored it without your permission? These are all moral questions, not technical ones...
It wasn’t designed, it’s just a very common metaphor about the perception of things rather than the way they came to life. It means that we should embrace it instead of trying to fix it.
> You now speak about entropy, which one?
Boltzmann. A system where information is preserved forever through the arrangement of energy states is highly improbable, so regardless of individual moral choices and the effort it will fall apart by the laws of nature.
Bow down to Unix children of Macintosh...
The whole post kept the same biblical style while describing why the Mac was conquered buy NexT.
A really great post that every once in a while I try to find on the internet.
Hard to say what is lost when it is unknown.
Is it possible to lose something you never knew?
But usually you realise after the fact learning that it was when you could have interacted with it.
Those things exist then through legends and the collective that still remembers, though that collective is very prone to just misremember.
If it's just lost I might still find it. I believe the post I'm looking for is lost. I don't know if it's gone. If it's gone I will never find it.
Could be it if it's lost there must be knowledge of it?
"Have you checked all the drawers?"
On HN we see every day interesting, first page content marked with a title representing its year, sometimes dating back a decade or more. Its age isn't apparently a detracting aspect to this audience if the content is still worth sharing.
And from the article the headline figure also doesn't represent irrelevant/undesirable content either but Wikipedia references, news articles, government pages, along with less unexpectedly ephemeral things like Twitter posts.
We're lucky there is archive.org but since it's not indexed like a regular search engine the only tether to old pages are still-live links found via regular search engines/sites (including HN). Essentially unless sites continue to exist that contain links to archived content the chances of future discovery becomes slim.
My stance is if you find something interesting that you expect is worthwhile sharing try saving it in the most convenient single file page format available to you (MHTML, SingleFile, PDF), to have your own copy. For MHTML at least it also saves the original page URL in its metadata. Saving to online archives is also great but admittedly higher friction (and can sometimes result in things like IP restrictions on archive.org even when saving just a handful of pages in a row, ime).
The core of the web pages on the internet are still there. It's just that the thick layer of commercial cr'app' websites built up on top are transient.
You don't update your server or database or runtime/framework/library? You'll get hacked and will drown in CVEs. You do try to do these updates? Have fun rewriting bits of your code, because the old version of a framework/library is no longer supported and there are breaking changes, which mean needing a partial rewrite.
Your best bet around that might be one of the relatively stable databases like SQLite, a micro-framework on the back end for a RESTful API, a simple solution for auth like basicauth/mTLS/... at the web server level in front of your API and then something without a toolchain on the front end, like jQuery. I mean this unironically, unless you want to maintain very few sites, then you probably have a bit more time on your hands.
Feels to me like the only content that can have any sort of longevity without constant investment of time is static sites - where updating your web server or moving to a different one is trivial and there are no write operations involved in most of the processes (maybe setup logrotate or just delete the logs occasionally).
But if the page needed to fetch data from an unmaintained API server that ran out of disk space, lost its DB network connection, got rebooted by a VPS provider, or any other issue, that site will probably never work again.
You don't get to choose if a third-party decides to rewrite an API interface, deprecate an entire library, etc.
You know that proprietary code has a cost because you pay for it. "Free" software is added to projects without much thought of what happens down the road.
I drag and drop a folder of PHP files underpinning a folder of Markdown file content, and voila: a website. Works online and offline (local server).
There was a really nice music school/venue/coffee shop in the town I used to live in. It shut down, and they took their Facebook and Instagram pages down. The only evidence it ever existed is in posts announcing events on other Facebook pages, in memories of people who went there, and on still surviving business listings that are likely to go away.
Aaron Swartz's website is still online more than 10 years after his suicide: http://www.aaronsw.com/
Someone has to keep paying the bills and renewing the domain. Still no HTTPS though. Does anyone know how this is handled?
This is either added since his death, and it's maintained by supporters, or it was there to begin with and one of the several people took over maintenance and funding.
But that still leaves the question of who. Maybe they want to remain anonymous.
I run a few daily word games. They're static sites that could continue forever as long as they have players, but sometimes I think about what would happen if I died or stopped paying attention. Domains would expire. Maybe some people could still play cached versions, and archive site versions would still exist.
I host them on github pages so if I exposed that url then that could probably exist for as long as github does - or breaks something that requires the owner to click a button to fix or something.
I could probably publish a free desktop version on various stores, and that might last as long as the store, or OS upgrades break it.
It would be great if there was a way to publish something such that it exists as long as anyone cares about it.
Can't you export your data from social media platforms? I haven't tried it, but did Google this for Facebook[1] and it looks as though you can?
[0] a now-deleted comment
[1] https://www.techlicious.com/how-to/how-to-download-all-your-...
However I am proud to say that you can still see my very first published web page from 1995 if you know the rather obscure url... http://admin.benwillies.com/ticker/
I wrote this page as a proof of concept for a friend of mine who was a financial consultant but unfortunately the humor was a turn-off instead of getting him excited about this new fangled thing called the Web.
I never made that mistake again.
This isn't a facebook thing, most facebook migration happened a decade ago. Instead this is companies closing, campaign websites shuttering, urls being changed, community events being over, and so on.
I know Archive.org archives a lot, but they can't archive everything, especially all the small personal websites.
Storage is expensive. Site owners don’t want to hold stuff forever, and the Archive cannot afford to hold everything forever either.
this seems extraordinarily high. since they cant do that estimation on all tweets ever sent it must be a bias of their sample
Google, Bing, Archive.org and others I contacted and had them remove my content.
Then I removed everything visible.
Very very happy I did, especially with the rise of AI.
Go back few years already and most links are dead.
Linkwarden is an open-source collaborative bookmark manager to collect, organize and preserve webpages:
The web we have today is just worst than pre-Google dominant era.
https://medium.com/luminasticity/fruit-stealing-scoundrel-ha...
and quite a lot of the content I used is no longer found, so all that remains is the word cloud but not the original poem.
It's a bit worse than the article details though because often the web page is there, or if not is available from IA, but much of the content was actually inside of comments and so forth.
(project was originally inspired because back then you couldn't go anywhere on the web without someone hitting you over the head with their This is Just to Say parodies)
Today there are over a billion. So those 38% amount to 22% of today's internet
Seems like a pretty good half-life!
There is a lot of valuable and interesting data on the Internet, but it is not visible. Certainly high quality, low profile blog that ended its development in 2015 will not be ranked high in Google.
Media platform, search engines monetize content. YouTube channels need to churn new content every week or so to stay relevant and to stay watchable.
Our society produces content, not quality, not products.
SEO can be gamed, it is impossible to create objective index of valuable content. Bad actors will hack the game, spam results, destroy quality to gain profit.
Google search engine most often connects users with media sites, with news sites, with the middle men. The more often not connect users with product directly. Write "search engine" in search query, you may not only find search a "search engines" but articles about "Best search engines in 2024", or "best SEO tricks to boost your page".
Google does not have any incentive to fix this. Search engines are dead tech. It will be replaced by chatbots in a few years. People will not search for content, content will be generated at wish.
Some time ago I have created my own domain repository with domain names: https://github.com/rumca-js/Internet-Places-Database
I wanted to find "wargames" related pages. It is quite impossible to find anything interesting concerning warhammer on the normie internet (not Facebook).
The second thing is I cannot find anything "amiga" related.
This solved this my initial problem. I have also found out that many interesting pages are gone. I think that Google directing our attention toward "content" broke good quality pages.
Right now I am using less and less google, because I use more and more my bookmark manager.
https://github.com/rumca-js/Django-link-archive
My solutions may not be as complex as common crawl, but they are enough for me. For now. I am still working on my program. It has been fun and interesting experience for me, and I learned a lot. About open graph protocol, about schema, about web scraping, etc. etc. Maybe this will inspire people to be more self sufficient, and more self-hostable.
In times of walled gardens we need more standard, and more open data to keep what remains of the old wild west of the Internet.
https://www.copyright.gov/fair-use/
more likely to find that nonprofit educational and noncommercial uses are fair.
less likely to support a claim of a fair use than using a factual work (such as a technical article or news item).
if the use employs only a small amount of copyrighted material, fair use is more likely.
And since code tends to be more idea than expression most of it can be considered to not fall under copyright after the application of the Abstraction, Filtration, Comparison doctrine.
Of course if you piss off an entity with a bunch of money to throw at lawyers it could be a bigger issue, regardless if you’re in the right, because defending yourself can rack up legal fees.
I do copyright/patent/trade secret inspections of source code for a living.
EDIT: Yet again, downvoted for just stating the truth... The irony is that 10 years ago and before the controversies around LLMs this comment would not have garnered negative attention because the forum was all for weaker copyrights... when copyright affected musicians instead of programmers' bottom lines... Sigh...
EDIT: I did a little bit of investigation and there are similar limits on copyrights in Germany known as "limitations on protected rights" that seem to carve out things like educational use and archiving, but I don't know anything about German law. I would find it surprising if most Western nations didn't have something similar to fair use unless there was an active interest in damaging the public's access to information.
Many legal systems don't know "fair use" and by default you have effectively zero rights to do anything with copyrighted materials without explicit permission.
The license will tell you what is allowed (and If its a standard one you can assume it is in accordance with the law)
We are maintaining 25y old urls is a bit ehh. cumbersome, and I sometime wonder if its worth it. Most of the traffic seem to come from bots and they do seem to learn some of the 301's. It seems to be good for SEO, etc.
Some users also gets redirected to the content on the new urls. It feels a bit like helping an elderly person over the street to where the shop is.
Anyway, I hope that bots and humans trust our services more.