Link rot and content drift are endemic to the web
theatlantic.com
theatlantic.com
The main crawler still seems to be heritrix3 (https://github.com/internetarchive/heritrix3), but there's a great little ecosystem with tools such as webrecorder and warcprox.
Still, I've read through the code of these tools and am feeling that they are failing in the face of the modern web with single page apps, mobile phone apps and walled gardens. Even newer iterations with browser automation are getting increasingly throttled and blocked and excluded from walled gardens.
Perhaps the time has come for a coordinated, decentralized but omnipresent approach to archival.
I keep thinking back to Jacob Applebaum's stance of "facebook and the other walled gardens are the real dark web."
If only there are a way to algorithmically tie a proof of work for a new cryptocurrency to archival of the internet in a way that wouldn't be easily gamed (by people archiving easy to access content or highly redundant archival of trivia).
There was a time when I was furious with the web going to hell and I investigated the possibility of "web without browsers" that started with making a WARC capture of page and putting pages through extensive filtering and classification before the user sees anything.
With interactive capturing you can push a button to indicate that a page is done "loading" but with automated capturing you can't really know that the page is done or that you got a good capture. That ended the project right there.
At the time I was most bothered by the slow load times of web pages and blaming this phenomenon:
https://www.sjsu.edu/faculty/watkins/samplemax4.htm
particularly that if you take the max of N random variables, the expectation value you get gets worse as N increases -- that is, the page isn't done loading until the slowest http request completes.
So I saw the "knowing when the page is done" problem as being particularly core, and it would be if the goal was to "win the race" against a conventional web browser.
If you were (say) preloading all the links submitted to hacker news you might be able to tolerate the system taking 5 minutes to process an incoming page. (See archive.is)
Today I've noticed that sites like Wired are giving up on complaining about my anti-track and ad-blocker and they just load the page partially which would drive me crazy if I was serious about debugging.
Maybe now that times have passed, people have died who posthumously admitted their preferences for white supremacy (and heavily bitcoin-supported that), and whole projects have been renamed, there can be a more inclusive community built around browserless web? For those who haven't followed, i'm refering to Woob (previously Weboob). I'd be interested in other people's feedback about that community lately, the ideas are great!
Could you expand on that? It seems a bit out of (anti-)left field, so to speak.
Nice to know about Weboob. No idea what the community was like but it's nice to know I'm not insane for thinking about stuff like this.
The article has the unstated assumption that eternal preservation of all writing ever is a net benefit. I think it's worth having a discussion on that point.
I suppose you meant "decrease"
The overall bent is a hand-wringing about link rot, which I thought we mostly got over a decade ago. The Internet is fundamentally ephemeral. If you see something you like, save it so you can repost it later. If you rely on someone else to keep in up indefinitely, you're being foolish.
Around the edges of that main discussion, the Atlantic also touches on censorship in all the wrong ways, re-iterating the too-common view that censorship is good as long as the good guys do it. They at least argue that this censorship should be transparent and censored works still accessible in some way, but they seem to not understand the nature of what they're talking about.
Censored works aren't censored to protect the public. They're censored to protect the rich and powerful. That's why "right to be forgotten" really exists. That's why Google and Youtube and Twitter and Facebook quash anything that goes against the accepted narrative in any given field. They aren't protecting the public from dangerous misinformation. No one gives a shit about the public. They're protecting the financial and political interests of some very powerful people.
Given this, talk of a "poison cabinet" only illustrates ignorance of the issue. The Memory Hole cannot be divorced from the censorship process, it's a core part of it. If people can still find the information in some form, it's not censored enough to make the people it threatened happy.
And this leads to the final point, which is that the real reason the web is "rotting" isn't link rot, it's censorship on the part of tech monopolies, due to their joined-at-the-hip relationship with every large corporation and industry imaginable due to advertising and other deals. The fact that links die doesn't matter much: you can just repost the material. The fact that links ARE ACTIVELY KILLED to suppress their information is a much more serious problem, and one that doesn't have an easy solution besides full breakup of the tech monopolies.
The misinformation area is somewhat stickier, but here's a decent example: if somebody decided to hurt you by spreading rumors (let's say that you watch CP) and spends time and money to get that rumor top of any SEO and forum thread, what's the right course of action? How good are you going to feel about using speech to counteract that when the result is a Google search giving your denial in spot 1 and the accusation in spot 2?
We all have to grapple with power and the ability to abuse it, but I don't think it's effective to say power is fundamentally wrong. The conversation is more nuanced than that and has to be viewed as systems with checks on power, which means specific design-thinking.
We have libel laws to address that.
GPs point is that Google, Facebook, et. al. are premptively censoring non-mainstream content just to protect themselves. They don't really care about the public.
To protect themselves from the public. Whether it's because consumers might take their business (and their data) elsewhere in disgust at what a particular platform is turning into, or because democratically elected lawmakers could start imposing sanctions or new regulations.
Companies are always looking out for themselves, that's a given. But that doesn't mean their actions are completely divorced from public opinion.
Did we? Should we?
Acknowledging the current state of affairs doesn't require accepting its flawed nature.
Imagine a world where gasp the BBC, NYTimes, etc. kept all versions of their articles available and online. Where the "pretty URL" shows the most recent version of a page, provides a permalink, and provides permalinks to all previous versions of a page.
I don't expect most sites to do this, but since someone else mentioned the BBC, I am targeting journalism as an example.
Not just that, but even due to politics, someone wanting their history "changed", or even worse reasons.
There were multiple reddit arguments recently, when someone posts a newsstory about something, discussion starts, two hours later the story is changed (without any footnote or old version available), and arguments start, because "there is noting in the article that says <what you said>", and no way for the original poster to prove, that it was, because noone even thought they'd need a screenshot.
This is a particularly good quote to sum up the article. The internet is not a repository of facts, it is a repository of facts, spam, junk, and things. Moreover, it is not the only repository of these.
Link rot happens. Content is subject to the will of the publisher to spend the time and/or money to continue to host it.
Depending on links to work eternally is a mistake. The problem is not the link rot, it is the bad assumption.
I know somebody who started a business that was successful for a while and then failed. Spammers got control of the domain and now it is full of ads for a dangerous diet drug.
What makes my blood boil is that it impugns the integrity of the founder who is a decent person who has nothing to do with that scam.
Back in the 1950s they put Wilhem Reich in jail, where he died. L. Ron Hubbard got the hint and left the country and when no country was safe he went to sea.
Today people like Dr. Oz run alt-health scams continuously and nobody seems to go to jail or even get a fine.
Abandoned formerly-popular domains create a kind of long-tail info-environmental impact, just like an abandoned warehouse can become a real-world hazard.
Maybe we need a digital superfund process.
We shut down the business an abandoned the domain. Someone registered the domain, created a similar-looking website by hand (recycling a lot of the text and images), and added spam. It even has my old company address.
This is a .st domain, about $35/yr. The web design work cost something too. More than I would have expected the link juice from a single website to be worth.
What we really need is some sort of DNS record or meta content we can add that tells search engines "this domain is being abandoned, destroy all link juice".
The original PageRank paper assumed that PageRank approximated the distribution of views on web pages assuming that people followed links at random.
If Google wanted to know what people are viewing today, they don't need to collect a link graph and do matrix math. They can measure it directly with Chrome, Google Analytics and data exhaust from the advertising platform.
Good stuff should be preserved, but it's not the Internet's job to somehow magically do it. It's OUR job, and the nature of digital information (DRM not withstanding) makes this easier than ever.
<and please skip the tired argument that I should just pay a subscription fee to avoid their ad crap - we are already paying them with our data. My data is worth far more to me than the value they provide for it. Plus - I won't give money to a company who forced this Faustian bargain on me.>
Cron job which runs a batch script at midnight which feeds youtube-dl some playlists URLS's which it downloads to a HDD. I also have nextcloud running which has access to the directory the videos are saved to, so I can easily share them if I want to.
In fifty years how many people will care? How many people should care because it would mean ignoring the huge volume of newer stuff? A hundred years? Two hundred?
I haven't even read or seen many of the existing cultural artifacts we have from past decades and centuries, what would I do if orders of magnitudes more of them had been preserved?
In fact, I'm incredibly greatful that I grew up before all the random shit that I threw out there as part of my youth was subjected to obsessive cataloging and archival efforts.
So while I get what you are saying that much of the pop culture videos do not hold long-term value (which is also questionable considering how many of us older folk still have collections of vinyl)... there absolutely is valuable content out there that deserves preservation.
And the same will be true of old vs contemporary sources in fifty years.
I don't expect my own collection of books - which includes some that are valuable to me primarily for nostalgia - to have much value past the death of myself and the rest of my generation. It might temporarily have a lot of monetary value near my death - when other copies have already been lost - but to someone born fifty years after me? What use would pulp fiction from the 80s be to many of them?
My parents and uncles are in a bit of disbelief of how little even I care about the Beatles already, after all.
We don't live dramatically longer or have dramatically larger memories than our ancestors, so things necessarily have to get lost and replaced by the new things that have been created since then.
Also, important to whom?
You might have a better understanding of the culture that produced them. You might appreciate a work of art that would otherwise not exist. We have graffiti from Pompeii, we know Ea-nasir sold cheap copper in ancient Ur 3700-odd years ago, but we've lost countless works of literature, music and film, some by the greatest masters of their age. What artifacts of culture survive the scouring sands of time is often a matter of happenstance, rather than quality.
Chances are almost everything our species has produced culturally, scientifically and artistically - the whole corpus of our knowledge output over the last century - is going to vanish within a generation or two anyway, simply because the digital foundation into which we've transferred so much of it is brittle and ephemeral. If we want to leave anything behind for future generations at all besides climate change, pollution and nuclear waste, we should save as much as possible rather than only what we consider to be relevant.
The idea is that a future person could freely deep dive through a rich well indexed history of media about whatever specifically interests them
I wish people would stop trying to assess the value of a given piece of media and just tag and archive the stuff.
For instance, high quality footage of live music from 100 years ago would be very interesting to some.
Our society is characterised by the constant generation and exchange of massive amounts of information. It's one of the things that sets us apart from previous generations. Preserving only a small subset of that data that we deem worthy or important will not allow future generations to fully understand today's society.
Plus smaller sites will start disappearing because of regulatory capture. It will not be possible to run a forum or similar site in few years.
Everyone’s assuming that data now stays on the internet forever because it’s so massive. It’s usually one or two people who keep the flame alive
Not only this, but many of the largest companies these days would never remove advertising even from a paid service. It's like cable TV. They wanna charge you and advertise at you for more money.
Download it? This continues to be not difficult for YouTube.
> I suspect that one day, they will make their ads unblockable by embedding them in the video files.
That's fine. I honestly wish they would because most of the hangs in YouTube I experience when the stream changes to an ad, and then changes back. If it was embedded in the video then the stream wouldn't be interrupted.
If I hate the ads that much one can edit them out after it's downloaded.
Fuck no. Do not do this. Use youtube-dl[0], and maintain local copies of anything useful you can find.
It skips sponsors (even from the youtuber itself), jingles, intro, etc.
It's awesome.
as an example:
"You don't have to trust me on this one, here's an article with [a bunch of data] | [*Archive link in case of link rot]"
from: https://kolemcrae.com/notebook/virtue.html
It's not perfect, but it helps reduce some of the issue.
Other than that solutions are incredibly hard to come by - you need institutions to preserve urls - through tech changes and the like, when they have very little incentive to do so. Eg. making sure they implement a redirect from the http to https sounds simple enough, but not everyone did it. Also if they switch CMSs and the like.
Note that you should also have a rule to save the link content locally, to avoid single-point-of-failure problems in the unlikely-but-catastrophic case that archive.org itself goes down. (Cf the attempts to attack them over their National Emergency Library programme last year.)
Had the web somehow been centralized (I have no idea what that would even look like), content still would not be archived, it would be constantly changed, and subject to censorship. Just like in a decentralized web, perhaps even more so.
Archiving costs lots of money (and costs keep growing if you only add and never take away), can be highly challenging (in the case of web apps or complex dependencies), whilst providing zero immediate reward for the organization carrying this heavy load. Not only is there no incentive, many couldn't even afford to if they wanted to.
And it gets worse still. Digital archiving means paying forever. Imagine paying a 100 years of electricity, hardware replacements, migrations. The entity (business, person) is long gone before that.
As a ridiculous example of this: Facebook has several very large idle content data centers. Mega scale buildings full of servers storing photos of Facebook users they haven't accessed in years, and likely never will again. Yet should a user do this, they expect the photo to still be there.
That's why I believe the problem should be addressed with more pragmatism. Focus on things of unquestionable long term value, and think of a good solution for this smaller scope.
I don't sit in the camp that everything digital must be preserved and that it's a disaster if it isn't. I try not to fight entropy in it's many manifestations. It's a shame when content disappears but I think it's also healthy to just accept it. We tend to only frame information disappearing in a negative light because we can always imagine a scenario where that information could have been valuable to someone, and that is a valid concern, I just don't think it's helpful to view it as the internet going into some downward rotting spiral and therefore every single 0 and 1 must be preserved.
The major problems of the internet seem to be almost entirely cultural currently.
But, for guarding against "published", supposedly static material disappearing, or changing silently, and for removing a short list of organizations from being responsible for preserving content, IPFS or something like it seems well-suited. Anyone who cares to preserve something can. Any change is noticeable.
IPFS has its flaws, but content addressing and immutability are powerful.
Such efforts deserve a guaranteed permanent home with permanent funding. Another example: hundreds of 'Old-time-radio' programs and early TV shows have only emerged and survived because of enthusiasts. If they disappeared from Youtube ....
It'd be good to set up a universal 'permalink' library system (UPLS) like the DOI system and its 'persistent interoperable identifiers'. This could (and, at least in part, must) be publicly-funded. Anyone who wishes could apply for a unique ID. Then, subject to a set of specs from some participating library ("yes, we'll host that"), they could package that content with metadata. Backups must be ensured. If eventually the content needs to move (or can't find a home), the ID (and metadata) remains.
Lots of books, magazines, newspapers do not survive because nobody wants to read them. We don't keep a super-archive of every piece of paper ever published, because we acknowledge that it's OK for things to die. Important things we keep, useless things we throw away. It's irrational to hoard garbage.
So really, all we need is to proactively decide what we do want to save, and save it as soon as it's published. Maybe also provide a "digital shelf" for people to keep their own copies, and a standard way to search for and distribute that information, like NNTP + Gnutella for e-books (along with some features to avoid being sued a la DMCA). The rest should be allowed to rot like fallen logs in the woods.
No man ever steps in the same river twice, for it's not the same river and he's not the same man
Heraclitus
https://en.wikipedia.org/wiki/Filler_text#:~:text=%22Now%20i....
# signifies an anchor
:~:text= signifies a text link
%22Now%20is%20the%20time%20for,21%20(1918). says show me the text between "Now is the time for" and "21 (1918)."
You only lose what you cling to
Gautama Buddha
Entropy is king. Eventually all information loses its coherency.
The invention of wood block printing 1300 years ago in the Tang was apparently specifically motivated by the desire to preserve and reproduce Buddhist sutras; the oldest surviving documents printed with movable type, from 900 years ago, are also Buddhist texts.
Of course the Tripitaka is not permanent; it will be lost some day. But you seem to be implicitly claiming that Buddhists do not apply effort to preserving information and in particular textual records, because they know that ultimately they will be lost. In fact, the truth is quite the opposite, and believing your implicit claim would require almost complete ignorance of Buddhism, printing technology, and South Asian classical studies.
Sometimes this helps them whitewash their screw-ups which have lead to widespread false beliefs. For example, after the UK government targetted and hit 100,000 Covid-19 tests in a day, the BBC ran an article falsely claiming Germany had achieved this a month earlier and linked it prominently on their news front page for about a month. A large proportion of the population probably saw this and now falsely believe it, it got brought up all the time as part of the narrative that the government's big "world-leading" achievements were just playing catch up badly, but it was memory-holed from the article in a rewrite and they used that as an excuse for not publishing any correction - so unless historians dig deep in third-party archives, they'd never understand where that belief came from. (Apparently a previous version of the article also wrongly claimed France was carrying out more Covid-19 tests due to mistaking their weekly numbers for daily one, according to a correction which disappeared from the article after a few days and only exists in the Internet Archive now. I haven't been able to find the original version of that claim.)
This is very true. Sometimes it takes 5-10 seconds to load the calendar view for an archived page, and another 5-10, or more, to load a snapshot.
They have a ton of data to manage with limited resources, but it still seems it should be possible to go faster than this. If there's just not enough budget for I/O, maybe they could offer a donate-for-data-dump option, where you can donate in exchange for loading data of interest (say, BBC archives) into a storage medium or query engine of one's choice, so one could do research at a much faster pace.
archive.org is much less annoying if one avoids using a bloated browser and Javascript to make HTTP requests
here is a more lightweight approach, not nearly as klunky/slow, IMO (393 bytes)
usage 1: (simple html page of all results)
echo https://www.theatlantic.com/article/619320/|1.sh >1.htm
firefox ./1.htm
usage 2: (retrieve last result)
echo https://www.theatlantic.com/article/619320/|1.sh 1 >2.htm
firefox ./2.htm
#!/bin/sh
read x0;
x1=web.archive.org;
curl -s "https://$x1/cdx/search/cdx?url=$x0&fl=timestamp,original" \
|case $# in :)
;;0)( printf "<h2> $x0</h2><ol><pre>\n";
sed -n "/[0-9]\{14\} [hf]/{s/\(.* \)\(.*\)/<li><a href=https:\/\/$x1\/web\/\1\/\2>\1<\/a>/;s/ </</;s/ //2p;}";
printf "</ol></pre><br>\n" )
;;1)curl -s $(sed -n -e "s>.*>https:/$x1/web/&>;s> >/>" -e \$p);
esac
haproxy + nc version
(965 bytes)maybe it is faster than curl, maybe not; you be the judge
#!/bin/sh
read x0;
x1=web.archive.org;
printf "defaults\ntimeout client 50000ms\ntimeout server 50000ms\ntimeout connect 50000ms
\nglobal\npidfile $HOME/1.pid\nfrontend f\nbind 127.0.0.21:80\ndefault_backend b
\nbackend b\nserver s ipv4@207.241.237.3:443 ssl ca-file /etc/ssl/certs/ca-certificates.crt\n" \
|exec haproxy -D -f /dev/stdin;
printf "GET /cdx/search/cdx?url=$x0&fl=timestamp,original HTTP/1.1\r\nHost:\40$x1 \
\r\nConnection: close\r\n\r\n"|exec nc -n 127.21 80 \
|case $# in :)
;;0)( printf "<h2> $x0</h2><ol><pre>\n";
sed -n "/[0-9]\{14\} [hf]/{s/\(.* \)\(.*\)/<li><a href=https:\/\/$x1\/web\/\1\/\2>\1<\/a>/;s/ </</;s/ //2p;}";
printf "</ol></pre><br>\n" )
;;1) printf "GET %s HTTP/1.1\r\nHost: $x1\r\nConnection: close\r\n\r\n" \
$(exec sed -n -e '/[0-9]\{14\} [hf]/{s/\(.* \)\(.*\)/\/web\/\1\/\2/;s/ //p;}'|exec sed -n \$p) \
|exec nc -vvn 127.21 80;
esac;
if [ -f 1.pid ];then kill -9 $(sed b 1.pid);exec rm 1.pid;ficurl version
#!/bin/sh
read x0;
x1=web.archive.org;
curl -s "https://$x1/cdx/search/cdx?url=$x0&fl=timestamp,original"|case $# in :)
;;0)( printf "<h2> $x0</h2><ol><pre>\n";
sed -n "/[0-9]\{14\} [hf]/{s/\(.* \)\(.*\)/<li><a href=https:\/\/$x1\/web\/\1\/\2>\1<\/a>/;s/ </</;s/ //2p;}";
printf "</ol></pre><br>\n" )
;;1)x=$(echo x|exec tr x '\002');y=$(echo y|exec tr y '\003');z=$(echo z|exec tr z '\036');
curl -s $(sed -n -e "s>.*>https:/$x1/web/&>;s> >/>" -e \$p)|exec tr -d '[\02\03\36]' \
|exec sed "s|<script src=\"//archive.org/includes/analytics.js?v=|$z$x 1|;s/<.-- End Wayback Rewrite JS Include -->/$y 1/;
s/<.-- BEGIN WAYBACK TOOLBAR INSERT -->/$z$x 2/;s/<.-- END WAYBACK TOOLBAR INSERT -->/$y 2/;" \
|exec tr '\036' '\012'|exec sed "/$x 1/,/$y 1/d;/$x 2/,/$y 2/d;";
esac
haproxy + nc version #!/bin/sh
read x0;
x1=web.archive.org;
printf "defaults\ntimeout client 50000ms\ntimeout server 50000ms\ntimeout connect 50000ms
\nglobal\npidfile $HOME/1.pid\nfrontend f\nbind 127.0.0.21:80\ndefault_backend b
\nbackend b\nserver s ipv4@207.241.237.3:443 ssl ca-file /etc/ssl/certs/ca-certificates.crt\n" \
|exec haproxy -D -f /dev/stdin;
printf "GET /cdx/search/cdx?url=$x0&fl=timestamp,original HTTP/1.1\r\nHost:\40$x1 \
\r\nConnection: close\r\n\r\n"|exec nc -n 127.21 80|case $# in :)
;;0)( printf "<h2> $x0</h2><ol><pre>\n";
sed -n "/[0-9]\{14\} [hf]/{s/\(.* \)\(.*\)/<li><a href=https:\/\/$x1\/web\/\1\/\2>\1<\/a>/;s/ </</;s/ //2p;}";
printf "</ol></pre><br>\n" )
;;1) x=$(echo x|exec tr x '\002');y=$(echo y|exec tr y '\003');z=$(echo z|exec tr z '\036');
printf "GET %s HTTP/1.1\r\nHost: ${x1}\r\nConnection: close\r\n\r\n" \
$(exec sed -n -e '/[0-9]\{14\} [hf]/{s/\(.* \)\(.*\)/\/web\/\1\/\2/;s/ //p;}'|exec sed -n \$p) \
|exec nc -vvn 127.21 80|exec tr -d '[\02\03\36]' \
|exec sed "s|<script src=\"//archive.org/includes/analytics.js?v=|$z$x 1|;s/<.-- End Wayback Rewrite JS Include -->/$y 1/;
s/<.-- BEGIN WAYBACK TOOLBAR INSERT -->/$z$x 2/;s/<.-- END WAYBACK TOOLBAR INSERT -->/$y 2/;" \
|exec tr '\036' '\012'|exec sed "/$x 1/,/$y 1/d;/$x 2/,/$y 2/d;"
esac;
if [ -f 1.pid ];then kill -9 $(sed b 1.pid);exec rm 1.pid;fiRe slowness, we're able to do a lot with a little, but there's always room for improvement. If you're interested in some of the specific infrastructural challenges, I did a presentation in February:
https://archive.org/details/jonah-edwards-presentation
and my colleague did a fantastic presentation detailing some of the internal workings of the Wayback Machine just last week:
https://www.4chan.org/robots.txt
Also, I let a domain of mine expire and the new domain owner (which just plastered ads) had a robots.txt that retroactively removed my “previously archived website” from the Wayback Machine.
news sites should really be mandated to keep previous versions of their newsstories with all the edits, especially the ones paid by taxes.
I expect future historical tooling will exist to solve exactly this problem. Assuming Archive.org and the like nabbed it, the evidence is all there for future generations to see.
This should show us that most of the web isn't worth preserving anyway, much like McDonald's burger wrappers aren't worth preserving like sacred artifacts. Most web and social media content is worth less than said greasy burger wrappers.
Better to post links to trusted sources and let people judge for themselves.
Whenever I take an unpopular stance I remind myself of Rick Sanchez's wise words, "Your boos mean nothing, I've seen what makes you cheer".
I'm glad I no longer pay for a TV license.
You notice it cos you are in tech, but the same happens in financial news, science and even sport. Go read a tech publication.
For shits and giggle I did once try to get a technical story on how to copy DVD's published - it got very heavily edited! http://news.bbc.co.uk/2/hi/science/nature/1987665.stm
(I'm a former + early BBC News website employee)
> Britons 'baffled over euro rate'
> Wireless internet arrives in China
> Mobile spam on the rise
Fascinating to see how much our problems have stayed the same, despite the changing context.
I hope this is considered 'archived' and not 'forgotten'.
And realize that the best place to preserve history is the Internet Archive' Wayback Machine.
Kind of the same way newspapers were never responsible for maintaining their archives, but librarians did on microfiche (remember that?).
But I'd take it farther.
First, the Internet Archive ought to have an official partnership with the Library of Congress and other national libraries across the world, that help provide funding. It shouldn't have to rely on private donations.
And second, it's time browsers integrated with it -- if content no longer exists there should be a built-in option to easily check Wayback Machine with a single click, and use a heuristic to show the most recent "good" version.
In other words, let the Wayback Machine be not just a, but the place for the Internet's history. Let's make it official.
(Also, Twitter posts are often linked, but usually the text is copied for archival purposes.)
My experience with Wikipedia, forums such as those, and Wayback Machine and arxiv.org make me think that people will do a lot of stuff basically for free and that by building communities, you don’t need extremely clever trustless incentive systems like Blockchain or major paywalls (although granted, the forum does have something like a paywall for unverified pre-public info) or massive platforms with multi billion dollar companies in order to disseminate information, analysis, news, etc. Best practices of web forums from the 2000s (active moderation, a sense of common purpose, expectations of non-toxicness, etc), are a really good solution.
What always baffled me is why forums didn't embrace technologies we had in the BBS days. Offline readers were the greatest thing since sliced bread - you could use whatever interface to a message board you liked, whatever editor you liked, etc. It was a lot easier to quickly scan through literally thousands of messages with a native, local client that relying on the constant ping pong between your client and a remote server.
Decentralized aggregation is what we really need. A combination of RSS and DNS. Ways to foster creation and discovery of hand curated lists like the original Yahoo - but thousands of them. No reliance on Google, Facebook, Twitter, etc. It's a nice dream anyway...
The bad design and low quality content is a symptom of the Internet's broken underlying economics. That's a human problem, not a tech problem.
Perhaps there will be many personal archives like mine that one day can be shared in a similar vein to copy parties.
We will need to treat the information we find online with its impermanence in mind (as authors, making things easy to copy, and consumers, copying stuff).
Perhaps it is this mindset that, when sufficiently prevalent, could make the internet more like a library again; weed out the garbage und curate the nuggets.
Btw I think archive.org is doing God's work but I don't believe any amount of coding and crawling will be able to save everything (nor should it). It can capture some raw data for (future AI?) historians to sift through though.
https://wiki.archiveteam.org/index.php/ArchiveTeam_Warrior
You can choose to help archive reddit, pastebin, URL shorteners and other ephemeral parts of the internet
https://wiki.archiveteam.org/index.php/Warrior_projects
I've also taken to updating the citations in Wikipedia articles with archive links.
The problem is that the foundations are shifting sands, and we need something that has significantly more integrity at the bottom layer, we can't just bolt URNs on as an afterthought. Some organizations are able to maintain persistent data over time, but it is in spite of the technology, not because of it.
I will also note that a world where it is possible to delete things is a world where individuals can be made to have written anything in the past. On the internet, at a certain point the past can be fabricated from whole cloth.
edit: and ironically, the issue is that this is because the internet wasn't actually academic enough in its original design.
[1] That's the best way to get a better title. https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
1. Failure of Lumen/ChillingEffect initiative to prevent bogus takedown requests. Currently, anybody is able to takedown any page on the Internet by sending bogus requests and then takedown any mention about who did it.
2. Google's failure to “organize the world’s information and make it universally accessible and useful”. As the author said: “no such transparent, academic competitive search engine exists in 2021“.
I see some relation between those two.
- even the best spa out there often barely shim the normal browser behavior. Yesterday my back button broke once more, in 2021. Infinite scroll don't let you pinpoint your position. User don't expect being able to copy / paste a link to take them to the content anymore.
- the ecosystem of url handling is fractured. This month I worked on a Django + React app, and my clients asked that it should be able to handle being hosted behind an arbitrary URL prefix if provided in the conf. Here are the things I had to tweak:
* adding the prefix to the proxy pass apache conf (yes, they are still using it);
* adding the prefix the react router conf for which most tutorials were outdated;
* adding the prefix in the js bundler conf as the base, and for the dev proxy;
* adding the prefix in the <base> element in the main template;
* making all urls and ajax calls relative to the <base>;
* making all the react router Link and history.push _absolute_ (took me a while to figure this out);
* serving the index.html as a template file from django, not nginx, to inject all that stuff according to the env var.
* hacking the build script to replace static files URLS with template place holders because the js bundler didn't have a hook for that (thanks sed);
And that's on top of the regular work of making urls in SPA works, which implies sync your backend and frontend URLS for pages and API. Who is going to do all that works? In fact, how many devs have the knowledge to do that? Pre-SPA, there would have been well documented 2 steps to do the same thing. The junior in the team could figure it out.- we had a ton of manure on top of our urls. AMP. Url shorteners. Tracking ID and redirections. Content wall. Captcha. Often several of them at the same time. If one of them break in the chain, goodbye URL.
- low code mean low skill devs, that never heard the mantra "cool url don't change". They don't even know they should care.
- some browsers just hide the URL. The users don't know what an url is anyway.
- apps don't care about deep linking. They could handle url fine, mind you. We have the tech for it. But it's not even on the radar of most devs. You don't address the content, you consume whatever pops up, so why bother ?
Plus, google is so good at finding the content you want out of the barely readable drunken mess of letters you feed it that most people don't type url anymore. People don't care about URL just like people don't care about bees dying, because it's too abstract to worry about.
<a href="https://example.com/some-article"
data-ipfs-warc="0xB45165ED3CD4 ...">
Some Article
</a>Now you write an article about a poster, and instead of a photo, there is an embedded instagram photo from the artist, who then removed their account, and the content is missing. Or a hotlink to the authors website, and it's missing there too.
Yes, if the book was destroyed too, all info about the poster was lost, but atleast the book didn't say "go to the corner of X and Y street, and hope it wasn't removed or destroyed by the weather".
A photo of an inscription is not the same as the inscription itself --- detail is lost in any translation.
For digital works, it is possible to faithfully reconstruct an original with full fidelity (if necessary, embed an emulated environment of the host, server, or network originally provisioning the work). But copyright claims make this a legal suicide maneuver, at least for any entity capable of being sued, or being sued effectively.
Note that numerous previous archivists have in fact been pirates or copyists, sometines under pre- or non-copyright regimes, but very often in direct rejection of copyright.
"But I'm the consumer of a service!"
--Ask the service provider to open source their work.
Not passing judgement on the decision to take away his posting privileges. But by suspending POTUS, everything he posted during his term in office is just... gone. Every hot link to anything he said, on any website, is broken.
This is an enormous loss to any historian of the era. He was using Twitter as his main microphone to speak to the world, and all that content is, while not lost lost, thoroughly and permanently scrambled.
It would have been better to just lock him out of the account, publish a statement that the @realDonaldTrump account is now permanently archived, and that any new account he tried to open will be suspended.
Stop being pessimistic
How they get implemented in solving this problem is the question
How do you hash content which is programmatically determined and changes on every page load?
How do you account for the same work in multiple versions, translations, or updates, strictly using hashes?
(Note that a chained hash, e.g., a git history, is not a strict use of hashes, though it most definitely does use hashes.)
Ethereum's Solidity is a great real world example. The instructions are compiled into a binary, and that binary can be addressed via its cryptographic digest.
Changes in state are easily represented as instructions.
>multiple versions, translations, or updates, strictly using hashes?
The same way a git repo does it now: Good software design.
>is not a strict use of hashes
Huh? I'm not sure what your trying to convey. Are you referring to the fact that git uses diffs between commits? At any commit, a repo may be re-hashed and the current state completely represented by a digest.
Of all things to worry about being represented as a digest, software seems to be one of the smallest concerns.
Merkle trees are useful.
They include hashes. They are not simply hashes.
Content-addressable storage robust against variations is another approach.
Didn't the original post say, "How they [hashes] get implemented in solving this problem is the question"?
Merkle trees implement cryptographic hashes.
>Git chain hashes.
Git chains hashes? Yes, of course, along with diffs.
How the permanent web will implement digests in various applications is the question. Whatever the answer is for link rot, the guaranteed solution will implement and depend upon cryptographic hashing algorithms. There are already many examples of this class of problem being solved by utilizing hashes.
It's simply that hashes, alone, do little, and that hash-free solutions might well exist as well.
Bald assertions of simple necessary and sufficient solutions are almost always mistaken.
More specifically, it's an on-demand publishing mechanism, relating to media-based publishing (books, records, physical video media) much in the same was as the electrical grid (a transmission mechanism) does to fuels (a storage function).
It would be nice (in at least some regards) if publishing content at a specific addressable URL were a promise to 1) eternally provide that resource and 2) never change it. But there's no way to guarantee that this will be the case.
In the past, the means for achieving archival of information was:
- To specifically record that information in some form. As with, say, Plato's memorialising of Socrates's dialgogues.
- To create multiple copies of those recordings, so that loss of any one instance doesn't mean total loss.
- To define a refererencing or indexing system such that individual works can be identified unambiguously. (Or more specifically, with an acceptable level of ambiguity.)
- To develop a means of agreement as to what the canonical or recongised version of a work is, or absent that, of identifying more canonical forms or lineages. (The historiography of the history of philosophy is an interesting sub-field, with a good introductory treatment in Peter Adamson's History of Philosophy Without Any Gaps podcast.)
URLs of and by themselves address these needs poorly. At best they point to a location at which a document may have been available at a point in time. A URL plus a time-range (which is what the Internet Archive's Wayback Machine effectively delivers) is a much better approximation to what is needed. (The notion of time-bounded identifiers generally seems a useful one, as might be applied also to domain and user names, for example.)
As I see it, the goal of the archival web needs to address a number of points which Zittrain addresses only very indirectly:
- Identification of what really should be archived. Right To Be Forgotten exists for very well-founded reasons, and a world without forgetting (or with very capricious forgetting) is one form of hell.
- True document-centric identification. I've been thinking about this for a while (recent discussion here: https://news.ycombinator.com/item?id=27455520), and something that is based on the actual contents while being resilient to mild changes seems most optimal. Checksums, not so much, tuples or ngrams, possibly warmer.
- A number of archival institutions. The Internet Archive is certainly amongst these. Other libraries, if at all possible spread across multiple institutions and jurisdictions would be preferable. IA have been working with numerous academic institutions and the US Library of Congress, though IA itself still carries most of the burden.
This isn't a new problem. In particular, each time there's been an explosion in some new form of publishing, there's been a scramble by archivists to keep up. Denis Diderot, 18th century encyclopaedist, has an awesome quote about the information explosion drowning his own generation (the encyclopaedia was his technical solution to that problem).[1] Numerous elements of what we now accept as standard bibliographic elements (titles, authors, tables of contents, indices, references, citations, page numbers, paragraphs, inter-word spaces, ...) were invented to answer specific needs, not always of archivists as a principle focus, though often providing benefits to them. Cataloguing and classification systems likewise.
I appreciate Zittrain's message. He's crying over a lost cause and looking backwards, not to the future.
________________________________
Notes:
1. Diderot: https://www.historyofinformation.com/detail.php?entryid=2877
Have they or have you stopped looking? I'd say the former.
Entire new genres of creative output - music, fiction, fandom, films, cosplay, hobbyist and enthusiast communities have been spawned by the modern web. It's never been more vibrant.
I'll never understand why people on Hacker News seem to believe the internet stopped evolving as an expressive space just because services replaced the need to design websites by hand. That's like believing literature ended once scribes were replaced by the printing press.
It used to be a lot like HN - discussions around links to articles. I wish there was a community with the feel of HN with the wide net of Reddit.
It feels like all social media is converging; Snapchat, Instagram, Facebook, Reddit, TikTok, Youtube, all an endless stream of ai-curated short videos that you can swipe through over and over.
It's likely that AI curation is the future because no humans can shift through the vast amounts of data and content being made. Things that can't keep up without curation or with just human curation have died or will die.
Good AI curation can bring you the exact content you're looking for, can but not will. I've seen it work times and times again, but I've also noticed that you have to be aware of the flaws of the tool to be able to use them or it gets really bad really quick.
You can't let the AI take control, if it derails to content you don't like you must know what it uses as a quality signal and give it a thumbs down, if it is intentionally derailed, you must stop using the platform.
TikTok recently released an update to their algorithm, it ruined my FYP and replaced my content with inane videos made by people nearby - hyperlocal garbage. The feedback mechanisms given no longer work, before that update they did.
I do think that even people being nostalgic here about the "old internet" should try and learn how to turn AI curation for their own advantage instead of just being sad and nostalgic.
The same is true for desktop, with with the Reddit Enhancement Suite browser extension. My Reddit has looked largely the same for nearly 10 years!
And as soon as the mainstream knows about it, and it no longer feels quirky and niche, it will be declared dead and abandoned anyway.
I've hit bad links before. Four out of five times, I can do a general search for the title of the document that should have been at the link or the quoted excerpt that the document I'm reading pulled from the link, and I get a clone of the document posted somewhere else.
I agree that it is necessary today, due to the sheer amount of useless sites that pop up on page 1 of the search. I wish Reddit invested more into making their internal site search better. If people did their searches directly on websites, Google would have an incentive to improve search results so it wasn't always the same 10-20 websites topping the list for nearly every query.
The second law of thermodynamics applies to the Internet as well.