Cool URIs Don't Change (1998)
w3.org
w3.org
If someone here is running any part of the Rust infra, please consider getting this redirect back.
*Eye twitch*
Otherwise we’d still be speaking proto-Indo European arguing about words that came from whatever language came before that.
Related:
Cool URIs Don't Change (1998) - https://news.ycombinator.com/item?id=29442818 - Dec 2021 (72 comments)
Cool URIs don't change (1998) - https://news.ycombinator.com/item?id=27537840 - June 2021 (140 comments)
Cool URIs Don't Change (1998) - https://news.ycombinator.com/item?id=23865484 - July 2020 (156 comments)
Cool URIs don't change (1998) - https://news.ycombinator.com/item?id=11712449 - May 2016 (122 comments)
Cool URIs don't change. - https://news.ycombinator.com/item?id=4154927 - June 2012 (84 comments)
Tim Berners-Lee: Cool URIs don't change (1998) - https://news.ycombinator.com/item?id=2492566 - April 2011 (26 comments)
Cool URIs Don't change - https://news.ycombinator.com/item?id=175199 - April 2008 (9 comments)
I would have agreed with you a year ago. But now, with AI, the entire web model may become less of a thing.
I thought the web would endure forever, despite the attacks from platforms and browser monopolizers.
Now it looks like social will soon be overrun by robots (maybe creepy ID verification will save it?), and the need to publish information will diminish as you can generate it and probably share the results on some future AI platform.
What does that even mean? Because it could credibly mean like 20 different things, some of which would clearly be nonsense, but you haven’t left us with a specific enough of an idea to engage with either way.
AI applications function with the web model. The web will to some extent endure forever, I doubt it'd ever go fully extinct. We may use AI tools more, but not for everything.
I've found the key, when you're on a forum and see something discussed more than once, is to change your thinking. Instead of getting annoyed, think great, a new cohort of people will get to talk about this important topic.
This kind of reframing works in many different areas of my life and the older I get, the more I have to apply it to avoid becoming a grumpy old curmudgeon.
(Links to the old discussions are great though, they just don't need to be followed up by complaining every time.)
Repetition may be seen as a visitor’s expense, but appreciating it creates a huge net profit over time, cause it helps a community to stay fit and alive.
> No man ever visits the same URI twice, for even if it is a cool URI which does not change, he is not the same man.
But the sheer amount of times where sites get redesigned and somehow every link to the site breaks at once is utterly ridiculous. There's zero reason to change your URLs every time you change CMS, and there's even less of a reason not to redirect the old URL format to the new one.
Yet somehow it happens constantly, especially on news sites which seem to love changing their site structure every few months or something. Sigh.
Are you willing to pay them for that work?
No?
There's your reason.
If we as end-users want URLs to not rot away, we need to put value on working URLs that convinces webmasters to put in the effort to maintain working URLs.
I think a lot of folks originally thought of websites like reference books. And some should still be thought of that way (Wikipedia, open source documentation, etc). I used to try very hard to keep all URLs functional during site redesigns. We accumulated thousands and thousands of redirects… most of which were never used.
I’ve since come to consider that a lot of sites are more like magazines: useful for a limited time span but not something that needs to live on your shelf forever.
1. Using URL schemes that come with the framework. Change of framework breaks everything.
2. Simply not caring. If the site is commercial, preserving rarely accessed parts for the sake of consistency is not important for them. PR carousel with big stock images and 1-3 sentence empty statements is the norm. You are not supposed to "use the site" you are supposed to come trough channels and campaigns that are temporary.
Everything is temporary. Archive what you care about if you really need it to last.
If you read the contemporaneous history from the early 1990s when the concrete of the web was still wet it should become obvious that it's worth a revisit of the fundamentals
For instance, DNS could include archival records or the URL schema could have an optional versioning parameter. Static snapshot could be built in to webservers and archival services would be as standard as AWS or a CDN; something all respectable organizations would have, like https.
These only sound nutty because it's not 1993 anymore and we've presumed these topics are closed.
We shouldn't presume that. There's lots of problems that can be fixed with the right effort and intentions. Especially because we no longer live in the era where having 1GB of storage means you have an impressively expensive array of disks.
Many unreasonable things then are practically free now
Farthest I got is that we probably should see two addresses where we currently see one in our URL bar: Locator and Identifier and whole web-related technology should revolve around this distinction with immutability in mind.
- On server side Locators should always respond with locations of Identifiers or other Locators. So, redirects. Caching headers makes sense here, denoting e.g. "don't even ask next five minutes". - Content served under Identifier should be immutable. So "HTTP 200" response always contains same data. Caching headers here makes no sense at all, since the response will always be the same (or error).
In practice, navigating to https://news.ycombinator.com/ (locator) should always result in something like HTTP 302 to https://news.ycombinator.com/?<timestamp> or any other identifier denoting unique state of the response page. Any dynamic page engine should first provide mechanism producing identifiers of any resulting response.
I feel there are some fundamental problems in this utopian concept (especially around non-anonymous access), but nevertheless would like to know if it could be viable at least in some subset current/past web.
I don’t think this can be solved by decentralized protocols. A lot of folks just won’t put in the effort. Quite a few companies already actively delete old content; there’s no way they are going to opt into web server software that prevents that.
Expectations are set, not interrogated. Let me give you an example
Companies and organizations with domains are expected to also be running mail on that domain.
Why? I can sit around and make up a bunch of reasons but none of them are given when that mail service is being set up, it's done out of expectation, just like how someone might pay $295,000.00 for the .com they want and wouldn't even pay $2.95 for the .me or .us
Are the .com keys closer together? Easier to type? Supported by more browsers? No.
There's mostly arbitrary social norms that get institutionalized.
They can go away. Having ftp service or a fax line, for instance, used to be one of them. Those weren't thrown into the trash for cost cutting reasons, the norms changed.
The question is where do we want these norms to go and what are we doing to encourage it?
This is how this could materialize - say there's an optional archival fee when registering a domain. Next search engines could prioritize domains that pay this fee under the logic that by doing it, the website owners are standing behind what they publish.
These types of schemes are pretty easy to fabricate - the point is the solutions are plentiful, it's all a matter of focus, effort and intentions.
Used it for those couple of blogs with zero issues, and now I can keep the originals around but inaccessible from the internet, so I don't have to worry about keeping them updated, aside from when I update php itself.
> I can keep the originals around but *inaccessible* from the internet
So it's only tangentially related to cool URIs not changing, since in this case the actual URIs did (apparently) go kaboom, this plugin just helps maintain the internal links within their local archive.
I like the flowers in your front garden. Please be cool and never change them. I like them there when I drive by from time to time.
No I won't just take a photo of them, or plant my own. I need you to maintain them, never landscape your property or be uncool. It doesn't bother me that it costs you money to maintain them, you should have planned for that before I had a chance to see your lovely flowers.
Please hand this note to the next home owner. Since I want the flowers to live forever and humans don't.
Sincerely, Idealism.
(tongue firmly in my cheek)
I have to go into the archives to look at the flowers, which means the archive person has to give me images from whenever they took the photo, where it might not be from the same year
I try to clear out my blogroll so people don't see broken links – but no one on that list owed me their permanent web presence.
I maintained 100+ domains over the years, and had to stop renewing most of them because my income drastically got reduced and I couldn't afford to renew them 'forever'. I was careful about which ones got nuked. They were typically sites that received very little traffic and had their heyday and fun in the sun, and there was no point in having them live on for perpetuity.
The few remaining ones I renew (sometimes 10 years in advance because ICANN) still get a lot of traffic and I regularly check the hosting setup to see if they're operating properly, and I check for 404s and downtime, or missing assets like images, JS, etc
A domain is something you commit to. If the project atrophies, you have to be willing to nuke it. But there will always be domains which are too good to nuke.
Well, that’s your problem, right there. Use one stable domain which you pay for, and use subdomains. That’s what the DNS is for.
Later I got tired of adding txt records, now my simple apps are like spa.bydav.in/weather or spa.bydav.in/radio spa.bydav.in/otp.html
I've brought up broken links in projects I contribute to on a few occasions and it seems like people basically don't care if there's even a tiny bit of extra work involved to fix it. Complaining about this and expecting people to maintain links themselves won't work.
URL-rules
URL-Rule 1: unique (1 URL == 1 resource, 1 resource == 1 URL)
URL-Rule 2: permanent (they do not change, no dependencies to anything)
URL-Rule 3: manageable (equals measurable, 1 logic per site section, no complicated exceptions, no exceptions)
URL-Rule 4: easily scalable logic
URL-Rule 5: short
URL-Rule 6: with a variation (partial) of the targeted phrase the page wants to get found for
URL-Rule 1 is more important than 1 to 6 combined, URL-Rule 2 is more important than 2 to 6 combines, … URL-Rule 5 and 6 are a trade-of. 6 is the least important. A truly search optimized URL must fulfill all URL-Rules.
My preferred URL structure is:
https://www.example.com/%short-namespace%/%unique-slug%
https:// – protocolwww – subdomain
example – brand
.com – general TLD or country TLD
%short-namespace% – one or two letter that identify the pagetype, no dependency to any site hierachy
%unique-slug% – only use a-z, 0-9, and – in the slug, no double — and no – or – at the end.
Only use “speaking slugs” if you have them under your total editorial control.
i.e.:
https://www.example.com/a/artikel-name
https://www.example.com/c/cool-list
https://www.example.com/p/12345 (does not fulfill the least important URL-Rule 6)
https://www.example.com/p/12345-prouct-nameFor example, suppose that you once wrote an article on AI, in, say, 2008, with the URL
https://www.example.com/a/current-developments-in-ai
When the subject, in this case AI, suddenly changes in later years, and you want to write an update, you’ll be forced to use this monstrosity of a URL:
https://www.example.com/a/current-developments-in-ai-2
Why not start with
https://www.example.com/a/2008/current-developments-in-ai
so that you can then create
https://www.example.com/a/2023/current-developments-in-ai
? (This will also benefit readers, similarly how HN likes having the year parenthesized in the submitted title.)
Metadata matters.
I've written ~2k posts and 1% ended up with suffixes like this (ex: https://www.jefftk.com/p/nomic-report-ii). And most of those were published in the same year as the original post anyway.
Then if you ever see the need to shuffle around your site, you can do redirects based on unique ID without needing to keep track of slugs or other metadata.
https://www.example.com/a/abcdefghi/current-davalopments-in-...
https://www.example.com/a/abcdefghi/current-developments-in-...
https://www.example.com/a/abcdefghi
are all the same.
manageable = measurable beeing able to seperate easily in analytics tools apple from pears, products from article pages from lispages.
while still having no dependencies in there
implies that rules 2 to 6 combined have negative value.
When I first learned Apache I thought that all URIs served by a webserver needed to be the same name as the file that they pointed to. Then I learned (embarassingly) much later that in HTTP the URI name can be anything at all and the content that it pointed to didn't even need to be a static file, that it could be dynamic.
I bet 50% of my bookmarks have become obsolete because the URLs weren't err cool.
So I'm developing a service that every day crawl your website and tests every single resource, be it regular links, images, CSS, fonts, and also external links over one has little control. And one of the features coming soon post-launch is getting notified if you change an URL and forgot to set a redirect from the old one, breaking SEO, bookmarks, and this very advice.
I'm starting a closed alpha this month and launch the next, so if you're interested in trying this out, send me an email.
EDIT: it's harder than it seems, because it is almost impossible to tell what is the new URL if you forgot to set up the redirect.
(Feel free to add earth.org.uk to your trial if you wish, email as here and/or on the site. I can share config with you.)
Do I put an ID that bad actors can crawl in ascending order? Do I hash the ID? Do I have a unique slug that requires thought to avoid clashing in the future? Does that add another DB lookup or are you looking up the ID anyway? Do you do /id-slug or /id/slug but if you move to Wordpress or off your stuck with /slug or something with no ID.
Maybe none of it matters but it seems like it does, especially if you consider it a change you expect to last the life of the domain. Every tech solution instead of rolling their own has their own unique way of handling it.
Heinous, even.
Destroyed, handle gracefully at different levels, and possibly in different ways across the site depending on editorial policy.
Things should generally not just disappear and leave URIs non-functional. Sometimes it does have to happen but those times should be true exceptions.
Wikipedia has descriptive URI's, presumably they change over time. I know some folks would prefer never breakage to anything else, but it seems like having a permanent URI with a UUID and a descriptive URI that could change over time, if, say, a building changes it's name:
https://en.wikipedia.org/wiki/Willis_Tower
https://en.wikipedia.org/wiki/Sears_Tower
One now redirects.
I feel like, yea, it's cool if your URI's never change, but as someone building a wiki, it's really always seemed a problem without a right answer. I'm honestly asking because the only way I know how to build my site with this feature is to just save the previous uri every time the descriptive once changes.
That's interesting to me, as a topic.
Some don't mind having that kind of problem, because it opens a door through which they can gesture at a veritable treasury of details. Details are interesting and very important to them. Perhaps they are even somehow the master of those particular details, for having considered them! Ah ha, great.
Others hate it, because they feel like they are springing an unfair trap on innocent passers-by who suggest a simple solution to what seems like a simple problem.
I can't say I can solve your problem, and for one it seems different now than it did before, in significant ways.
But having walked into it, I can tell you I know this kind of situation! I hope you enjoy puzzling it out, whether through others' ideas or your own well-calibrated wiki-developer brain.
Wikis typically keep a history, so they can show a “this topic has been removed” page that still lets you access the history. Or if the topic was merely moved, trigger a redirect. A redirect is perfectly fine, it means that links using the old URL continue to work.
I learned this the hard way when Wikipedia was updated to include some (iirc) dumb password advice (or maybe it was about hashing) and I nearly included that in a customer report because my snippet from a previous report used the normal en.wikipedia.org/whatever form instead of en.wikipedia.org/whatever?oldid=123.
Wish they wouldn't call it old ID: the latest revision is also referred to as such, and this just looks silly, like why is my consultant linking to a stale ID? Call it page revision or just ID or something... this discourages use of them but it's a core feature.
Also, Github links that die all. the. damn. time. Please use the permalink option when copying a file link: this will include the commit hash in the URL and so the link will also work if the file was moved, branch renamed, contents changed, etc.
I guess I weirdly find descriptive links valuable. I could easily do both, but I think just recording old links is probably an option in the long run. I just think people should be able to click on a URL and know where they're probably going.
This still doesn't explain what to do when DB entries are destroyed.
https://de.m.wikipedia.org/w/?title=Valheim&oldid=231146307 page title is still in the link, in addition to the revision ID
https://github.com/git/git/blob/d15644fe0226af7ffc874572d968... also says exactly where you're going, it just replaces the branch master or whatever with the commit. You still see the repository and file path.
Permalinks don't have to be obfuscated. The downside I would rather mention is that the link doesn't update to newer revisions if improvements are made. It's not like an LTS version where fixes are still applied but no changes (I guess because every change is supposed to be an improvement to the topic, but yeah, not always).
Cool. Now what’s an URI?
URIs != URLs
URLs == URIs