Short URLs: Why and How
sive.rs
sive.rs
A simple example is:
... where you can find "items":
https://www.kozubik.com/items/
... and one thing inside of "items" is an article on NDS emitters:
https://www.kozubik.com/items/nds/
... which contains supporting multimedia objects:
https://www.kozubik.com/items/nds/images/
Not only do you know where you are in the tree, but it is also discoverable: by using index.html default pages and your web servers directory indexing code, you can have a very flexible resource that is easy to navigate.
Have either:
example.com/<article_title>
or, if categories exist:
example.com/<category_name>/<article_name>
Now there are other reasons for hierarchal urls mentioned in the thread, like providing better semantic meaning when representing content that is already hierarchal. Though I find that that is rarely the case. For example, you might think to represent blog posts like myblog.com/articles/<title>. But you could also do it like myblog.com/<article year>/title (eg myblog.com/2017/why-choose-short-urls). I've seen both formats, and both are valid to a certain degree. So choosing one is a bit arbitrary. But if you allow both formats, that just pushes the ambiguity onto the user. For example if the user wants to bookmark the article, which url do they bookmark?
From my experience, data follows graph structures, so trying to force a hierarchal structure and doing things like encoding paths in urls, feels too arbitrary and adds unnecessary ambiguity.
Yes, and they should live by their choice and implement whatever is implied by their semantics.
URLs also have their own affordances and I expect to be able to act upon them.
If you've got a CMS, I suppose you could do both: store each article in its hierarchical place, but also give each article (or each article you consider important enough to share, but I would hope that's all of them) a short name through which it can easily be remembered and shared.
https://sive.rs/short-urls-why-and-how
would be much more valuable than
I do this on my own blog. e.g. a post titled "Mr. Robot Hides Data on Audio Disks, And So Can You! (Season 3 Spoilers)" has the URL https://jszym.com/blog/mr_robot_steganography/ and not the WordPress style https://jszym.com/blog/mr_robot_hides_data_audio_disks_so_ca...
You can't win, haha
I think I prefer a middleground. I like Sivers' short URLs but I'd go with a word rather than just letters. Letters seems like a step too far for my preferences, looks untidy.
In real reality I'd probably use both methods. Long SEO-delicious URL for the actual article, short redirect to it, best of both worlds. Use whichever is appropriate for the medium (preferring to link online to the non-redirect for less future hassle)
Regarding (a),I personally like the view that the content lives in a "flat world" on top of which we collate different structures to organize/filter a same set of contents. In that worldview, web users entry-points can be more than directory-listings. A great inspiration is how Wikipedia offers a way to find which articles use a given picture: each picture acts like a "category", the same way that "recent-changes" is another filter on the same "flat world" of articles.
However, what is immensely difficult is to standardize these in a world where de-facto implementations burgeon, flourish, and eventually becoming out of touch (e.g., site-maps, RSS, OpenGraph). Hence we are stuck with very limited but very generic tools (b) for which rules like "directory-listing on the slash-separator" or "generate a JSON of the whole site connections to display as an interactive graph" (which I do on my personal blog) merely are local work around which require a bit of duck-tape to work.
Don't believe me? Twitter, for example, has short, meaningful URIs for user's status pages, and long, meaningless-to-users URIs for tweets. And in the sea of tweets there is no hierarchy in their naming.
You: How does Twitter's lack of hierarchy demonstrate that hierarchy doesn't help anyone?
How did you read into my comment that I said that "hierarchy doesn't help anyone"?
"You can remember them"
Well, no. You might remember a few. But you are only going to remember so much. This has a very limited benefit. "You can tell someone"
Yeah, but no. I barely tell people a domain. I'll just send them a link on the spot. Again, very limited benefit. "They look nicer."
Ok, but not a strong use-case. In general usage the display name of a link does not need to reflect the URL of the link. "They remove the middle-man" ... "encourage people to copy and paste" ...
I'm just gonna say this is plain false as a generalization. I've worked on URL sharing in publishing for almost 2 decades (yikes).It's built on a little bit of truth. On mobile devices all have built in share functions to text and spread based on various apps installed.
Again, very limited impact and effect.
"They’re enough."
This is my favorite! He talks about unique combinations but this goes against his original point of making them memorable. He should of at least talked about potential dictionary level word combinations... sighWhy not?
> but they unavoidably have leaked and will continue to leak into it
It seems like the opposite has happened. URIs were the default navigation UX, but we are slowly pushing them away.
Because if you have enough of them for internal purposes then they can't have meaning to users.
> It seems like the opposite has happened. URIs were the default navigation UX, but we are slowly pushing them away.
URIs were the default entry point for navigation, and they still are. That's one way in which they leak, but these are intentional so maybe best not called "leaks". There's also the URIs that leak in other ways, like in the browser status bar, and so on.
However it is very limited benefit to the extent that the larger argument/conclusion (shorter URLs are better) is wrong.
It ignores many valid points raised here, like URLs that include structure (be it date-based, or taxonomy, or hierarchy) provide immense value in many cases.
I agree that URLs should be humane, but typeable from memory is not the primary objective. Having people beable to read the URL and understand something about where it will go is. Also, unless you know your domain will only ever host one type of content, then having no hierarchy will lock you in. https://sive.rs/blog/on-short-urls is plenty short, but so much more valuable to any humans using the url.
If you don't agree that URLs need to be humane, then none of it matters and sure, use a UUID for every page.
Example https://oin.am/rr/
> but my C drive crashed and I lost the bookmarks.
There's still a Delicious market.
https://web.archive.org/web/20080906141421/http://blog.delic...
>> "So why did we switch to delicious.com? We’ve seen a zillion different confusions and misspellings of “del.icio.us” over the years (for example, “de.licio.us”, “del.icio.us.com”, and “del.licio.us”), so moving to delicious.com will make it easier for people to find the site and share it with their friends. Of course the old del.icio.us domain and all its URLs will continue to work. Also note that the domain change requires a new login cookie, which is why everyone has to log in again."
https://latonas.com/blog/web-masters-episode-9-joshua-schach...
But beware, single letter local parts are not universally supported by web sites.
Microsoft accounts with those mail addresses are possible (they used to be impossible), but recently I stumbled upon two other web sites that didn't like single letter local parts.
I don't use the short oin.am as my primary (which is oinam.com), I have it as a alias domain, allowing me to quickly say/hand-write my email to someone and continue the conversation from my main domain -- say, ß@oin.am.
Are there many providers who support one, but not the other?
It's quite a convenient way to have site specific addresses while still only having a single mailbox to manage.
I sign up for accounts with emails like apple@corl.in and that lets me easily filter and organize my inbox.
But when I have to tell some customer service rep my email they get very confused.
Almost every email host that supports custom domains supports receiving email with unlimited wild-cards. Is there something I'm missing?
Example: https://www.fastmail.com/pricing/ ($3/month)
I'm also fairly sure that they don't let you reply as the address the catch-all was placed on, which is an important feature.
If you know any that do, that'd be great.
The point still stands though, if you have a mail provider that supports custom domains there usually isn't a limit on the catch all aliases you can add (or a very high one). In any case it's far from being "prohibitively" expensive.
> I'm also fairly sure that they don't let you reply as the address the catch-all was placed on
That is definitely possible with Fastmail as that's how I use it. In their case I believe that feature is called "Alias". If you reply to an email you just select with alias you want to reply from.
I know this is not available any more so I've done the same with Cloudflare Email Routing which lets you set up a catch-all and is (still) free.
When changing ISP, I gave ispname@myname.fr and the dude got confused and couldn't understand. He wanted to call his supervisor to ask if I was allowed to do that.
I checked every year on the expiry date to see if they renewed. Every year they did.
One year I forgot to check and now the domain is parked and for sale for $288k.
But for me, {firstname}.{country}, is enough because it's pretty cool to be the only {firstname} in {country} who decided to register the domain. :)
Yeah, maybe if you're SEOing for a content farm bot, but if you have actual content, that you and other people care about, add the fsking date.
At the same time I also always find articles that have an "Updated at" timestamp of a few days ago, I guess that's somehow done automatically to gain some recency points for SEO, not sure if that or removing the timestamp is more annoying.
I hate this trend, one theory I heard was that it helps the content appear "ever-green".
Another theory is that it helps hide inactivity of content updates; like 100s of articles will be produced in a short span of time and then nothing for months or even years.
But yes, I hate it.
e.g. from 2003: http://www.paulgraham.com/hundred.html
most recent post in 2022: http://www.paulgraham.com/heresy.html
I'm not sure I would want to have all my writing in a single directory, but I guess you can't argue with 20 years of longevity ... on both sivers' site and PG's.
Also... http?
One example of what can be done is to cause the users to DDoS a third party.
But I admit it's a valid reason to have https.
Hackernews is a global website last time I checked so even if you don't face issues like that, many other users might.
Not sure why you're talking about HN though, the website in question was not HN.
By using http you are assuming your "here" (where ever that is) is the same as where your readers are.
And if your "here" is in the US then hotel WiFi providers in the US are well known to have done MiTM to insert ads.
The man-in-the-middle attack potential other replies have mentioned is possible, but my reason is far simpler. Defaulting to https for everything removes the cognitive load of having to decide whether to trust a website and pushes users to believe everything should be secure by default. The environmental impact is, in my opinion, worth it.
Cool URIs don't change. Not even their styling.
It solves multiple problems for me.
1. All articles are sorted automatically.
2. I can visually see how many drafts I have (drafts are 00000-, published at 00nnn-)
Quick, what’s the URL of this discussion? Will you remember it? Does that have any impact on your usage of HN?
The vast, vast majority of website sessions start with a click. When people bother to type things in, they type them into search bars (reminder that in most browsers, the “address bar” is also the search bar).
So be clear what you want from your URL structure. If you want search traffic, you’d do better to organize it for Google (hierarchical topic-based) than making it short.
If you want to be able to say it or print it for easy typing, a domain name is going to be best. Catch it and redirect it to the target page.
Even “domain.com/word” is going to be hard for most people to remember. If they remember it at all, they will probably type in “domain word” and let Google figure it out for them.
Emphasis on generally.
URIs do leak into the UI/UX. It's unavoidable. Purists hate it. But it is a fact of reality.
/blog/2022/05/08/short-urls-why-and-how.html
I can tell it's a blog post, I can tell when it was written, and I can see enough of the title that I can remember if I already read it. The ".html" doesn't add anything though.Personally, I find putting the title of the post (but not the date) in the path is a good compromise:
/short-urls-why-and-how /posts/2022/modelling-workflows-with-finite-state-machines-in-dotnet/
This way people can still see the approximate date, while also ruling out the small chance in the future I'd want a post with the same title again (can't think of why, but I didn't want to rule out the possibility).Can you? I can remember some domains, but I can't recall many full URLs that I remember, except perhaps some that I use almost daily. Certainly not a blog post that I'll read only once.
> You can tell someone. You can even say it out loud!
If the other person is going to write down the URL, I'm not sure there's a difference here. Giving someone your email address over the phone is already a painful experience no matter what. Always have to make sure the dots are in the right place. And if the first point is wrong — "You can remember them." — then making it a little shorter over the phone isn't all that advantageous, because they're going to write it down or type it anyway, not memorize it.
URLs on sites with a ton of content often have a regularity to them. For example you might want to have different types of articles, say rooms, people and towels. For each article type you might want different affordances that usually have overlap, like an index, an input form, an activity feed...
URL paths are a good way to encode that kind of regularity.
Of course on a blog of a single person it’s whatever.
Similar: The variants of „Just write HTML!“ posts that come up every once in a while.
But I like the essence of the post.
You will get more creative in choosing titles.
- Instant links with low latency no matter where your visitors are in the world
- Robust and reliable links that will never go down, ever
- Protection against network and application layer attacks
- Unlimited tracked clicks (no cap on clicks/month)
The service allows the creation of links using a hierarchical naming structure, as mentioned in these comments like example.com/folder/test. It also auto-generates an SSL cert for every custom domain added, for HTTPS-secure links.
I'd love any feedback! The service is largely free and can be found at https://www.pxl.to
Those who care about SEO use long slugs for the extra juice. Those who don’t can use IDs or whatever works for them.
Shareable URLs - still need to share the domain, and most of us (including sive.rs) don’t have a particularly memorable one.
A simple, generic way could be a hash of the original URL with a compact representation:
http://example.com/some/path/to/stuff?filter1=asdf&filter2=90#positionX
http://example.com/s/Wq_2
/s stands for "short" and the ID is the hash representation. If you use a-zA-Z0-9 plus some URL safe characters like _ it should be reasonably short even for large sites. The CMS or whatever software doesn't have to to implement something, because it is mostly just a generic URL shortener running on the same host. <link rel=short href="https://youtu.be/abcdef"/>
Then copying the URL could copy the short link instead.That being said I like when the URL is readable. The YouTube example works because everyone knows that YouTube is a video. Presumably for non-video links they could use something like /user/foobar to differentiate from their "main product". But for blog posts I would much rather see something like /short-urls rather than /su as the reader gets nothing from the first one.
<meta property="shortlink" content="https://bit.ly/shortlink">
Just because it's not a standard doesn't mean you can't use it. This is sort of futureproofing incase it does become a standard.It looks what I described is very existing to the spec published here: https://microformats.org/wiki/rel-shortlink
I prefer things like example.com/cats when possible vs technically shorter but harder to remember things from bit.ly or whatever.
So short urls like this don't make sense if they are not explicit and you don't have a corresponding lookup table
In OP's case, /short-urls is succinct yet it hints at the content.
It’s just a nginx with redirects on pages))
I got a couple of them, but I end up using p0l.co, like p0l.co/55, also just a redirect list basically, on GitHub with netlify.
try_files $uri.html
<FilesMatch "^[^.]+$">
ForceType text/html
</FilesMatch>
Which I don't think is that bad, compared to RewriteRule's.try_files $uri.html
Though personally I would do as the other comment said, keep it hi.html and use try_files or mod_rewrite to handle it. I mean, ask me to set up something like this and I probably would never have though of doing it the way in the article. It just seems weird to me.
Seems perhaps similar to try_files for nginx that a sibling commenter referenced?
try_files {uri}.html {uri}
Done. (You can also add a redirect if you want to enforce that ONLY the non-.html version is used.)I was hearing Instagram influencers say "tap the link" all the time
so bought a domain https://tapthe.link which creates a url which can take viewers directly to the app via deep linking instead of a simple redirect,
super helpful for YouTubers and marketers and it's free for creators.
tapthelink is easy to remember than say a short url which says xlps.com or something
and don't do hierarchical URLs also
this are my battle-proven URL rules. whenever i violated them and changed the priority I regretted it.
URL-rules
URL-Rule 1: unique (1 URL == 1 resource, 1 resource == 1 URL)
URL-Rule 2: permanent (they do not change, no dependencies to anything)
URL-Rule 3: manageable (equals measurable, 1 logic per site section, no exceptions)
URL-Rule 4: easily scalable logic
URL-Rule 5: short
URL-Rule 6: with a variation (partial) of the targeted phrase
URL-Rule 1 is more important than 1 to 6 combined, URL-Rule 2 is more important than 2 to 6 combines, … URL-Rule 5 and 6 are a trade-of. 6 is the least important.the URL yoursite.com/short violated rule number 3. lets say you have list pages, category pages, tag pages, database generated pages, post pages .... now say you change your listpages, in design, in how they look, you want to know if they now work better (for whatever metrik you care about or not), well good luck with that. either you create a regex of hell for GA or whatever tool you are using or just a spreadsheet.
or you do example.com/l/short for list pages, example.cm/a/super-short for article pages, now easy to measure and manage.
also dont do example.com/category-name/article as all hierarchies are imperfect and change over time (there is no permanent or perfect hierarchy over time) so any kind of hierarchy violates URL rule number 2.
use %namespaces%, these are one or two letter words which identify the pagetype (listpage, article page) (note: pagetype defined as "share the same or very similar process on how the page gets created, changing the template or process of the pagetype, changes all pages under this pagetype)
i usually end up with
https://www.example.com/%pagetype namespace%/%permanent identifier i have under total control%/ or
https://www.example.com/%country or language namespace%/%pageytpe namespace%/%permanent identifier i have under total control%/ for multimarket and/or multilanguage webproperties
you can download my book for free here if you like https://gumroad.com/l/understanding-seo/hacker-news I go in a bit more detail there.
1. Public Suffix Domain (costs money to register).
2. Subdomain (255 char limit in domains)
3. URL (generally URLs aren't too long)
4. Content (basically free to stuff keywords into).
You don't know what you don't know. Discovery is about leveraging common heuristics. The more context information that's available, the easier it gets to get to the answer. Giving the least amount of information runs counter to that.
A URL is a reference, an identifier, that uniquely references a resource. A search engine essentially captures a ton of context information that you can leverage to get to a set of relevant references. An identifier can be a meaningless string of characters - e.g. a UUID - and as long as it's accompanied with context information, you can get to the resource it references.
Conversely, if you capture context in the URL itself - meaningful words, dates, authors,... - you're actually providing a breadcrumb trail for visitors - people and robots - to follow, leveraging common heuristics they might use to get to that web resource directly. Might, because discovery is always a process of making educated guesses and following cowpaths to get to the right answer.
So, no, making URL's shorter isn't necessarily advantageous.
In the same vain: long passwords using commonly, easily to remember words instead of an unintelligible string of 16 characters.
Tangentially, I also have ambiguous feelings over the widespread use of URL shorteners. Partly because they act as a middle man providing brittle URL's that can - and will - break in the long run. Partly because they hide a ton of potential context information that might be captured in the original URL.
> You can tell someone. You can even say it out loud! Whether answering an email or talking to someone on the phone, I can say, “Go to sive.rs/ff for my talk about the first follower.” or “My newest book is at sive.rs/h.” I do this often, so having memorable URLs saves me a lot of searching.
Again, context matters. This might work if you want to highlight specific content - e.g. a marketing page for your book - but it certainly doesn't work all the time for all your content. Is that blogpost from 2007 really that important that you need to be able to "say it out loud" to someone on the phone today at any moment?
Short URL's come at a cost. What's the trade off you're making here?
> They look nicer. They’re aesthetic. They show care.
I don't care. Really. I don't. I care about readability and accessibility. Sure enough, URL's with tons of non-nonsensical query parameters are a blight, but this has more to do with readability then "aesthetics". It's an URL, not poetry.
> They remove the middle-man. With long URLs, people use those ugly social share buttons that promote (and further entrench) harmful social media sites, and add visual clutter to your site. Short URLs encourage people to copy and paste the URL directly, which lets them share it anywhere, instead of only the sites for which you have a share button.
Or maybe the answer here is to avoid using social share buttons on your website at all?
> They’re enough. Using 36 characters (a-z and 0-9): 4-character URLs give you 1,679,616 (36⁴) unique combinations.You don’t need more than that.
Well, how about "sive.rs/qxfa" or "sive.rs/kxig" or "sive.rs/ddiz"? It's an argument that's in direct contradiction with the author's first argument. Again, heuristics matter. Readability matters. A big chunk of those 1.6 million odd unique combinations aren't usable of the bat because they are simply unintelligible strings of characters without meaning.
I absolutely will not. At least with known URL shorteners I know that’s what I’m getting and can inspect further. When I get links like this I assume I’m getting spam or malware. Almost all of the time it’s spam.
URLs, shortened or not, can send you anywhere. You don't know where they'll take you (unless you control them).