Examples of Great URL Design (2023)
blog.jim-nielsen.com
blog.jim-nielsen.com
https://web.archive.org/web/20030810201315/http://mpt.phrase...
He was on the search for his ultimate blogging system, where this "cruft-free" URL structure should be used:
https://web.archive.org/web/20051107103030/http://mpt.phrase...
I could have sworn there was a changeset in which Matt Mullenweg was implementing those cruft-free URLs in his new fork called Wordpress, but trying for google for something with "Wordpress" from the early 2000s is basically impossible in 2024.
Update: I found this: https://ma.tt/2004/08/mike-on-uris/
I currently work with an API which does a bit of content negotiation using the Accept header, so clients can request data in various formats - application/json for a snapshot, text/event-stream for an updating feed, or text/html for an interactive dashboard. I wish it didn't. I wish we'd just used file extensions. Trivial to use in a browser or via curl, trivial to implement on either side.
E. g. `/books` - looks at the `Accept` header. `/books.json` - sets the `Accept` header to `application/json`. `/books.xml` - `application/xml`, and so on.
I think that refers mostly to the .php and .asp of the time. Those don't tell a thing to the user.
But nobody wants webpage URL's that randomly end in .php, .htm, .html, .aspx, and so forth. That's just noise that is both gibberish and entirely irrelevant to the user.
But I agree about .php, .aspx and other extensions that are telling something about the server side. That’s irrelevant for the user.
it's _kind of_ relevant, if it weren't for the fact that the absence of any extension implies .html >99% of the time
Also, "you're not going to publish two essays with the same title" feels false. If you write 1,000 pieces and use short titles and tend to write about the same subjects, it feels extremely likely that you'll wind up repeating titles.
If we're talking about blogs/news, they don't ever get almost entirely rewritten. The original publication is the only date that matters, and it matters a lot.
If we're talking about evergreen content like documentation, then of course you don't put dates in the URL. A small "last updated" on the page itself is appropriate there.
Unfortunately, this isn't the case. It should be the case IMHO, but it (currently) isn't. The SEO/marketing people nowadays (ab)use popular pages for the search rankings and update them regularly to keep the content fresh and highly ranked (since search engines give much preference to new content).
Also, even for strict blogs/news, it's not unusual for a particular post to be a draft for many months before publishing. Most serious blog will fix the date to match publish date, but that isn't what happens by default especially in Wordpress (which is the most important platform for blogs).
Others seem to think just day and month is fine, as if the year isn't the most significant part. And if both numbers are <=12 then you have to go and find out what locale the author formats their dates in...
https://www.dahosek.com/the-big-countdown/
https://www.dahosek.com/the-big-countdown-2/
⋮
https://www.dahosek.com/the-big-countdown-11/
Alas, the default URL scheme in Wordpress doesn’t include the date.
Disambiguation is one thing, but as a reader, I really like having the date indicated in the URL for informational purporses. It's very helpful.
I use that for static files on my blog and it’s worked great for 20+ years:
https://static.simonwillison.net/static/2024/mlx-whisper-gpu...
https://static.simonwillison.net/static/2003/getElementsBySe...
https://web.archive.org/web/20020603092331/http://www.kottke...
http://scripting.com/2001/09/11.html (Every paragraph is in effect an "entry")
Then in the early 2000s blogging resulted in a style with longer articles instead of paragraphs. In the middle 2000s a "retro" style begun with far shorter and differentiated entries, the so-called tumblelog:
https://kottke.org/05/10/tumblelogs
The original Tumblr may have been inspired by this, if not just the name.
And then Twitter and other social media arrived on the scene and ate everything. :/
https://duckduckgo.com/?q=wordpress&t=h_&df=2000-01-01..2005...
I'm guessing you were looking for:
https://wordpress.org/news/2004/01/cruft-free-uris-in-wp-10/
I assume Google supports something similar, but I've stopped using it.
Poorly.
I have a blog so old I titled it an "online diary." It pre-dates search engines, so they tend to date the diary entries (blog posts) based on first crawl. Which means lot of the dates presented by the search engines are off by several years.
notion.so/:account/Current-Name-of-Page-:pageid
where the name changes if the page is renamed, but the redirect works, as the page ID is unchanged. In fact, one can just use notion.so/:account/:pageid
and gets redirected to the right page, or even notion.so/:account/Anything-else-:pageid
works too...This is very handy in my use cases, when various Notion data is extracted into another tool, reassembled, and then needed to have a link to the original page. I don't need to worry about the page's name, or how that name gets converted into the URL, or any race conditions....
The page hierarchy is then just within the navigaton, not in the URL, so moved pages continue to work too (even if this looks like a flatter hierarchy than it really is).
I'm sure there are plenty of drawbacks, but I've found it an interesting, pragmatic solution.
Kinda smart.
Also; taking it from the end means you only need to parse the string as an offset from the end. It can make load balancing much faster in theory.
The fact that so many sites do this (including "normie" news sites) shows that site designers clearly believe users want and expect "informational"/"denormalized" URLs, rather than /?id=123
The better way to implement this is to serve a 301 redirect if the words in the URL don't match the expected ones, that avoids trickery and also removes the risk of the same page being accidentally indexed as duplicate content.
So, if Paul Graham hosted his site via Notion (he doesn't), I could link someone to `https://paulgraham.com/Why-I-Hate-Hackernews-be2839f0-e145-4...` and it would show my (fake) page on PG's domain.
What do you do when the translated slug happens to be the same in multiple languages? I ended up still having the country code in the slug.
That being said, if the URL has a language code (/en/), it would be good to change the language code manually, and still end up on the right page. Sometimes the language switcher is really well hidden. However a visible language switcher is even better.
I use that ALL the time. I can navigate straight to any issue by typing a URL. I can switch to the “actions” view for a repo by adding /actions. I can see the file I’m looking at in a branch by editing the URL and swapping “main” for the branch name.
All available via the UI as well, but I interact with GitHub so often that the tiny efficiency boost I get from navigating by URLs really starts to add up.
I also trust them not to break links, based on their track record. My notes and blog posts and even my source code are full of links to issues or code snippets on GitHub.
/issues/
/issues/:id
/pulls/
/pull/:id /issues/
/issues/:id
/pulls/
/pulls/:author
/pull/:idIn rails, if you don't spend effort to undo, you get URLs that match your database setup in a RESTful way. For me, that's far to tight coupling, and problematic in other ways.
But the result is that without thinking about it, without a second of time spent on designing an URL structure, you get a very nice, consistent and clean one for free.
https://github.com/torvalds/linux/commit/{hash}
I also like that the "id" which is an ASIN (Amazon Standard Identification Number) which is a superset of all ISBNs. This means you can just enter any book ISBN directly into the browser and end up on the right page (at least historically) instead of having to search for it.
Maybe on some sites it's fine but I feel like link previews need more modularity on size or something. Perbaps even configurable both host and client
The information density of the preview has to match, if not exceed, that of the context in which it is provided. Most of the times it's a placeholder image with less text than the URL itself.
>I think this is actually a really great design in that the name of the product can always be in the URL
he said "before". you could accomplish your goal putting the slug "after". he's making the point that having a place after which you can harmlessly delete the rest of the url is better than having embedded NOPs surrounded by identifying information (not that the average user will ever edit any url, but there is still merit in what he said that you missed, and he's not disagreeing with you)
Stack Overflow, in contrast, puts the ID first, followed by an often very long slug. Which seems to be the more common pattern generally, as far as I can tell.
I do wonder what their rationale was/is.
Yeah, when I used Amazon I found this incredibly annoying. When I wanted to share a link, I'd have to spend a few minutes figuring out how much of that stuff I could remove and testing the resulting URL before sharing it. A relatively minor irritation, but an irritation nonetheless.
I don’t super like repurposing country names like that. Including .io.
It feels this disregards the actual meaning of the extension while ignoring some very legal consequences to be under Indian legal system instead of EU or US.
Format is reuters.com/:category/:headline:date
which is all you need to know what you're clicking on. For example, I don't need to describe this link in order for its contents - and its time-relevance - to be understood:
https://www.reuters.com/world/us-navys-newest-air-to-air-missile-could-tilt-balance-south-china-sea-2024-08-14/
Edit: they are a bit long, though, I supposeURL-rules
URL-Rule 1: unique (1 URL == 1 resource, 1 resource == 1 URL)
URL-Rule 2: permanent (they do not change, no dependencies to anything)
URL-Rule 3: manageable (equals measurable, 1 logic per site section, no complicated exceptions, no exceptions)
URL-Rule 4: easily scalable logic
URL-Rule 5: short
URL-Rule 6: with a variation (partial) of the targeted phrase
URL-Rule 1 is more important than 1 to 6 combined,
URL-Rule 2 is more important than 2 to 6 combined,
URL-Rule 3 is more important than 3 to 6 combined,
URL-Rule 4 is more important than 4 to 6 combined.
URL-Rule 5 and 6 are a trade-off. 6 is the least important.
A truly search optimized URL must fulfill all URL-Rules.
My preferred URL structure is:
https://www.example.com/%short-namespace%/%unique-slug%
https:// – protocol
www – subdomain
example – brand
.com – general TLD or country TLD
%short-namespace% – one or two letters that identify the page type, no dependency to any site hierarchy
%unique-slug% – only use a-z, 0-9, and – in the slug, no double — and no – or – at the end. Only use “speaking slugs” if you have them under your total editorial control.
i.e.:
https://www.example.com/a/artikel-name
https://www.example.com/c/cool-list
https://www.example.com/p/12345 (does not fulfill the least important URL-Rule 6)
They make URLs unnecessarily long, often forcing people to use URL shorteners -- completely defeating the purpose.
They get awkward when the author changes the title. Other commenters mentioned some tricks to get around this issue, but all involve redirects. Cools URLs shouldn't change in the first place.
They don't copy cleanly if you use nonalphanumeric characters, as in nearly every language other than English.
Virtually nobody just looks at a URL these days anyway, with all the search engines, cute thumbnails, and OpenGraph metadata that provide a glimpse of the actual content for you before you even click on it. This is doubly true in the non-English-speaking parts of the world where a slug in a shared URL is often just a jumble of %HEX.
Hand-picked words in URLs are fine, e.g. /about/me. I'm only talking about autogenerated slugs for user-submitted content above.
I haven't noticed this ever being an issue.
At least in Firefox, non-ASCII characters will show as-is in the URL bar, but in the copied URL they will be properly encoded.
One alleged justification for slugs is that they allow people to guess what the URL is about without actually visiting it. This usually happens in forums, comments, text messages, and other places where people can only use plain text. But non-ASCII slugs look undecipherable in exactly such places.
See all these pages here: https://kde.org/for/
e.g. RJF01 vs ab0fhct99fh2h4fqi2fj9
* have a separate table that maps slugs to IDs, allowing many-to-one relationship, because content's title will be updated, and you don't want to break old links.
* long slugs will get truncated by users. A zero-cost way to recover from that is `select id where slug >= ? order by slug limit 1`
* in either case don't forget to redirect to the canonical URL, so that people can't create duplicate or misleading URLs on your site.
`my.domain/post/123/my-great-slug-that-is-pretty-long-but-doesnt-matter/comment/456`
vs
`my.domain/post/123/comment/456`
Stackoverflow makes the slug completely optional but you have the choice of only accepting foo and bar in your example
Not to mention something similar can be done to any url, e.g. #whatever-you-want or ?_=whatever-you-want
eg. series/9876545678 and series/098767890 get treated differently and the analytics get difficult to merge. But really they're the same page just hydrated with different data.
Should've used query params, eg series?id=9876545678
Third-party component have to coexist with existing site navigation logic, so generally you can't safely add URL-based configuration to such a component.
Fortunately, configuration can now be stored in fragment directives in order to hide this from normal site routing. e.g.
https://example.com/page#routing-info:~:additional-routing-info-for-third-party-component
With fragment directives, location.href and location.hash exclude the additional content in the hash after :~:This is used in Transcend Consent Management for configuring parameters to debug and simulate various privacy experiences[1].
1. https://docs.transcend.io/docs/consent-management/reference/...
The idea was really cool, but from talking to people at HP at the time, the implementation was apparently a complete nightmare done with an insane number of rewrites. It was sort of a hit and miss if the thing you typed in after /go/ would actually take you to the correct location, if any.
firstna.me/lastname/
firstna.me/lastname/aboutThe projects formatted like: https://there.oughta.be/a/wifi-game-boy-cartridge
Do path parameters get to have / in their values?
Let’s say you have a link shortener service and want to allow users to define shortcuts like /mypath/:rest where rest is appended to example.com/
Now you’re in a very interesting position when it comes to resolving URLs.
Curious to hear folks with experience in this
, is encoded as %2C
/ is encoded as %2F
For instance, in query parameters, spaces are encoded as '+'. But what if '+' is also a valid character in your domain? You then need to disambiguate i.e between "name?foo+bar" meaning "foo bar" or "foo+bar". Which one is the user actually referring to?
In our case, we ended up needing users to send the name in the body, and now we have to manage multiple encoding protocols (url, queryparam, the body...).
A user might have entities named "foo bar", "foo%20bar", and "foo%2520bar". Sometimes, mistakes happened because users forgot to double encode or they used the wrong protocol. As this names were used in URL, query parameters, and the body, and each has its own.
As I mentioned, with clear constraints and rules, we can accomplish anything we need, it can get complex. My takeaway from this project is to limit the valid characters and make it simple for everyone.
Don't ask how I know.
https://stackoverflow.com/questions/16245767
Responds with a 301 and a location header value of `/questions/16245767/creating-a-blob-from-a-base64-string-in-javascript`
The downside is that it's a massive pain to explain to people verbally!
> stackoverflow.com/questions/16245767/how-to-bake-a-cake
Fortunately for SO the fake slug is not preserved and redirects to the real one (so e.g. stackoverflow.com/questions/16245767/motheficker is not served from their site), much to the chagrin those of us with childish sense of humor who some 25 years ago enjoyed dynamically generated nonsense like:
https://web.archive.org/web/20031007123544/http://john.isgay...
https://sentido-labs.com/en/library/201904240732/Xanadu%20Hy...
Something like Xanadu tumblers.
Personally I'd avoid using id's and use a 32-bit hash of the URL which is more or less as performant as a straight id lookup. I usually went with murmurhash.
mysite.com/en-us/some-page mysite.com/en-ca/some-page
You can 301 redirect some locale to your "base" URL if you want.
mysite.com/en-us/some-page > mysite.com/some-page
But don't stress too much. Google doesn't really care about URL content any more. People on phones don't care what your URL says. It's at most desktop users, and devs.
Don't stress localizing your URLs...
mysite.com/fr-ca/some-page is just as good as mysite.com/fr-ca/une-page... and the former is a lot easier to tie into email marketing variables.
Just keep your sitemaps in the localized folder.
mysite.com/sitemap.xml... just a link to the various localized sitemaps.
mysite.com/en-us/sitemap.xml etc.
By keeping sitemaps in a localized folder, it'll make it a lot easier for yourself as you go to register your site with each market's locale.
If you just have to localize URLs... consider doing what Amazon does and just tie the URL to an ID.
https://www.amazon.com/Moen-One-Handle-Bathroom-Deckplate-84...
the above is the same as this... https://www.amazon.com/dp/B0CFYPTKF8
And you can put anything you want in the URL string, it just matches on the ID.
https://www.amazon.com/literally-whatever-you-want-here/dp/B...
“We use the words in a URL as a very very lightweight factor. And from what I recall this is primarily something that we would take into account when we haven’t had access to the content yet… [but] as soon as we’ve crawled and indexed the content there then we have a lot more information. And then that’s something where essentially if the URL is in German or in Japanese or in English it’s pretty much the same thing.”
- John Mueller, Google Search Advocate.
I changed it to home.arpa, but I forgot why