Cool URIs Don't Change (1998)
w3.org
w3.org
I don't get Microsoft. They're huge. They hire a lot of people. Their products are kludges.
All documentation on Help Books was archived, for instance. It's been 7 years since they've seen an update and they now contain inaccuracies – but there are no other official guides. Check out that UI: https://developer.apple.com/library/archive/documentation/Ca...
This is a technology that is still used. Nearly all of Apple's own apps have Help Books, including new ones like Shortcuts. Yet they have absolutely no official documentation on using that technology.
Off-topic, but damn - tune down that candyness a bit and it looks much better and cleaner than what's there today.
URL indexing discipline > number of site URLs
(There was no CMS, every page was hand-written PHP. And to be frank, maintenance was FAR simpler than the SPA frameworks I work with today.)Hashes aside, allowing linking to a book by it's ISBN doesn't seem to exist either as far as I am aware, at least not without using Wikipedia's or books.google.com's services.
https://www.amazon.com/exec/obidos/ASIN/<asin id>
Say what you will about Amazon (and Jeff Bezos), but I don't think they've broken a URL to any product of theirs ever.Having public edit history + permalinks to specific immutable revisions of product pages would be nice, but I could see how they’re not incentivized to add it because they don’t want ppl demanding an older price / feature that got edited out later on.
If what you want is a library and a persistent namespace, you'll need to create institutions which enforce those. Collective behaviour on its own won't deliver, and chastisement won't help.
(I'd fought this fight for a few decades. I was wrong. I admit it.)
Preservation for infinity is competing with current imperatives. The future virtually always loses that fight.
It's all just the Golden Rule in the end; but the Golden Rule needs an accompaniment of knowledge about what struggles people tend to encounter in the world—what invisible problems you might be introducing for others, that you won't notice because they haven't happened to you yet.
"Clicking on links to stuff you needed only to find them broken" is one such struggle; and so "not breaking your own URLs, such that, under the veil of ignorance, you might encounter fewer broken links in the world" is one such corollary to the Golden Rule.
Keep in mind that when this was written, the Web had been in general release for about 7 years. The rant itself was a response to the emergent phenomenon that URIs were not static and unchanging. The Web as a whole was a small fraction of its present size --- the online population was (roughly) 100x smaller, and it looks as if the number of Internet domains has grown by about the same (1.3 million ~1997 vs. > 140 million in 2019Q3, growing by about 1.5 million per year). The total number of websites in 2021 depends on what and how you count, but is around 200 million active and 1.7 billion total.
https://www.nic.funet.fi/index/FUNET/history/internet/en/kas...
https://makeawebsitehub.com/how-many-domains-are-there/
https://websitesetup.org/news/how-many-websites-are-there/
And we've got thirty years of experience telling us that the mean life of a URL is on the order of months, not decades.
If your goal is stable and preserved URLs and references, you're gonna need another plan, 'coz this one? It ain't workin' sunshine.
What's good, in this case, is to provide a mechanism for archival, preferably multiple, and a means of searching that archive to find specific content of interest.
Personally I believe yes, because there are still those that benefit in the interim. Compare that to not bothering to fight at all in the first place.
There's also performing maintenance to defer the inevitable. I'm generally in favour of that.
But if you're getting exercised and emotional over something that simply and repeatedly proves not to work ... and there's some specific goal for that behaviour ... which can be more readily achieved by other means ... then I'd strongly suggest bailing on that fight.
TBL got some things right about the Web. He had a specific purpose in mind. The Web has (IMO very much for the worse) moved beyond that.
As I see it:
- The goal behind the advice is to create a long-term useful addressable archive of online content.
- The method advised ... has failed in this. Spectacularly.
- Other methods exist. They are being used, and work reasonably well.
- At some point you cut your losses.
That said, I'm cutting mine in this thread. Cheers.
A rewrite of a site does not rewrite all the content. It is still there, the database might have been migrated but all the information is still there and the conversion process has everything it needs to do the final step of preserving one ID, or one string, that trivially can map an old URL to a new one that contains the same actual content.
Sure, some intermediate pages might get lost but that is not something that is particularly valuable anyway and not something one usually links directly to. Don't let perfect get in the way of good enough.
The foundations, in a word, are profoundly unstable. The entire ediface has little robustness.
If you want preservation, find a better mechanism. URIs ain't it.
And rewrites are the most common of them and also being the easiest to actually have control over, and most can't even be bothered with that.
July 17, 2020, 387 points, 156 comments https://news.ycombinator.com/item?id=23865484
May 17, 2016, 297 points, 122 comments https://news.ycombinator.com/item?id=11712449
June 25, 2012, 187 points, 84 comments https://news.ycombinator.com/item?id=4154927
April 28, 2011, 115 points, 26 comments https://news.ycombinator.com/item?id=2492566
April 28, 2008, 33 points, 9 comments https://news.ycombinator.com/item?id=175199
(and a few more that didn't take off)
Note that dang will post these as well. He's got an automated tool to generate the lists, which ... would be nice to share if it's shareable.
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
(I email HN fairly regularly, mostly brief suggestions/fixes/issues on posts or threads, occasionally longer.)
A bit on dang's link tool here: https://news.ycombinator.com/item?id=28436784 (that links to further discussion). This describes the tool: https://news.ycombinator.com/item?id=26158300
And on the why:
After the creation date, putting any information in the name is asking for trouble one way or another.
Clearly these suggestions predate SEO.Personally about the only thing that has worked for me has been UUID/SHA/random ID links (awful for humans, but it's relatively easy to migrate a database) or hand-maintaining a list of all pages hosted, and hand-checking them on changes. Neither of which is a Good Solution™ imo: one's human-unfriendly, and one's impossible to scale, has a high failure rate, and rarely survives migrating between humans.
At the organizational level: my company has a dedicated “old urls” section in our urls.py routes files where we redirect old URLs to their new locations. Then on critical projects we also have a unit test that appends all URLs ever used to a tracked file in CI and checks that they still resolve or redirect on new deployments. Any 404 for a legacy URL is considered a release-blocking bug.
Here is a relevant Long Bet that I think about often (only has one year left to go!) https://longbets.org/601/ "The original URL for this prediction (www.longbets.org/601) will no longer be available in eleven years."
IME, in general, URIs for well-known sites (likely to survive) are more likely to keep working. Lately I've noticed many problems with sites that changed all spaces to underscores.
Relatedly, why does Firefox reload the page when duplicating a tab? If the page is a mediocre single page app, losing state essentially means being sent to the front page.
[1] https://www.morganclaypool.com/doi/pdfplus/10.2200/S00481ED1...
"Never do anything until you're willing to commit to it forever" is not a philosophy I'm willing to embrace for my own stuff, thanks. Bizarre how blithely people toss this out there. Follow the logic further: don't rent a domain name until you have enough money in a trust to pay for its renewals in perpetuity!
> Think of the URI space as an abstract space, perfectly organized. Then, make a mapping onto whatever reality you actually use to implement it. Then, tell your server. You can even write bits of your server to make it just right.
Oh, well if it's capable of implementing something abstract, I'm sure that means there will never be any problems. (See: the history of taxonomy and library science)
While I concede that the ability to retrieve the previous version of a page by visiting the old URL (provided anybody actually still has that old content) might come in handy sometimes, I posit that in the majority of cases people will want to visit the current version of a page by default. Even more so, I as the author of my homepage will want the internal page navigation to always point to the latest version of each page, too.
So then you need an additional translation layer for transforming an "always give me the latest version of this resource"-style link into an "this particular version of this resource" IPFS link (I gather IPNS is supposed to fill that role?), which will then suffer from the same problem as URLs do today.
But having unchanged documents move to new locations on the same domain without a redirect but a 404 is just utter unforgivable failure. Or silently deleted documents, also an uncool nuisance.
Both happen a lot. That's what comes to my mind, when I read the initial quote.
Let the webserver redirect the uri without version to the most recent one. Problem solved. Remember, redirects are valid, logical uris.
Because conventional URLs don't care about updates to a page's contents, this means the common case of only updating pages requires no additional configuration. I only need to invest additional time setting up redirects on the rare occasion that I actually do re-arrange file names etc.
IPFS URLs on the other hand are revision-specific, so they break (or rather actually get out of date and eventually break if at some point in the future nobody has a copy of that version cached/pinned somewhere any more) as soon as I fix even one tiny little typo, so I need to set up some sort of URL mapping service right from the start.
https://metrics.torproject.org/hidserv-rend-relayed-cells.pn...
The weakness is always the people.
Back to 2019 levels of bandwidth? I feel like I may be misreading that graph, but I'm more curious about what bandwidth suddenly spiked the last two years more so than the drop back down.
At any rate, I always understood the main point of Tor was providing an overlay network for accessing the internet and maintaining secure anonymity as much as possible, with its own internal network being more of a happy side effect.
I don't think IPFS would be as quick to kill compatibility with a vulnerable hashing algo compared to Tor since they're not aiming for security and anonymity as primary goals.
Some of the largest companies on the planet are actively opposed to this concept. If you care about this kind of thing champion it from within your own organization.
Pinterest is a great example of how organizations put business interests ahead of building an accessible web.