Trailing slashes on URLs: Contentious or settled?
zachleat.com
zachleat.com
I am guilty myself, because we need to do things that match expectation of users, but this problem simply disappears if you stop doing magic.
So, the real question is, what "magic" do you prefer?
My implementation was, when given a trailing slash, check if the recource is a directory (as this is what the trailing slash signifies) and use the convention of loading the index.html file at that directory if it is. If it's not a directory, then it's a 404.
If there's no trailing slash, try to load the file with the equivalent extension to the Accept content-type header. So, if the client wants JSON, you try /foo.json for the path /foo. If client requests HTML, try /foo.html.
Urls lost their 1:1 mapping to physical files a long time ago. Nowadays they are mostly just arbitrary strings, although well-crafted ones may reflect the hierarchy of content in their structure.
Says who?
Nearly every web server in existence supports static assets which map 1-to-1 to the file system.
> This is one of those "made up problems" as it only exists because you insist in doing magic, i.e. mapping /foo to either /foo/index.html or /foo.html.
If the underlying resource (web server) is Linux, then there is no magic. In Linux, everything is a file, including directories (maybe in Unix too, I don't know, I only work with Linux). So e.g. `/var/www/foo` is the same resource as `/var/www/foo/`. But without the context of the underlying filesystem being available to query that (e.g. with the `file` command) there is no way to know that `/foo` refers to a directory.Stumbled across this in Go yesterday
Compare a relativeUrl ref from 'http://localhost/1.0' vs. 'http://localhost/1.0/'
Use trailing slashes please.
"Because your CMS is rewriting out trailing slashes. Make it stop that, and then you'll be in the relative path you think your in".
The internet uncertainty principle: It's not DNS, it's trailing slashes. (apologies to Herr Heisenberg :)
[1] https://developer.mozilla.org/en-US/docs/Web/HTML/Element/ba...
I have no idea what Vercel, Render, and Azure Static Web Apps even are, but you should obviously be aware of their limitations when using them to host HTML. I'm sure there are plenty of other behaviours that differ between them and there's hardly any universal truth to be found in the decisions some big corporations once made.
This article seems more about the default slug-to-content resolution algorithm from a bunch of cloud providers than it is about trailing slashes.
You can see /1.0 as a file and /1.0/ as a directory, so the relative URLs make perfect sense either way.
return a string consisting of the reference's path component
appended to all but the last segment of the base URI's path (i.e.,
excluding any characters after the right-most "/" in the base URI
path, or excluding the entire base URI path if it does not contain
any "/" characters
https://datatracker.ietf.org/doc/html/rfc3986#section-5.2.3 <div dir="rtl">
http://example.com/example/
</div>
This displays as if it were the string: /http://example.com/example
It's awful. Much better if everything was configured to leave off the trailing slash.I would not have expected that URLs are detected inside rtl context in the first place.
Users will occasionally add trailing slashes, and will occasionally remove trailing slashes.
So as long as $PATH and $PATH/ map to the same outcome, and redirection from the ‘wrong’ one to the ‘right’ one to the other uses a 307/308 to allow non-GET methods to redirect, then everything will always work out okay.
Varying from that recipe in any regard is a source of pain and trauma in every complicated ingress scenario I’ve worked with for twenty years. (Dispatch methods, regular expressions, exact string matches, all of them.)
This doesn’t match my experience at all.
In the context of Linux and Windows file system stuff, it’s quite the opposite: all normalisation techniques that I can recall having encountered have eliminated trailing slashes, and it’s not unknown to treat a trailing slash specially in some way (e.g. rsync, or more tenuously Vim’s 'directory' option), and historically a lot of Windows stuff would choke on a trailing backslash (or on forward slashes at all), though it’s exceedingly rare now.
Of URLs, hmm… I can’t think of ever having encountered the addition or subtraction of a slash, except for the normalisation of URLs with no path (https://example.com → https://example.com/). Query string parameters, sure, but the path, never.
—⁂—
> and redirection from the ‘wrong’ one to the ‘right’ one to the other uses a 307/308 to allow non-GET methods to redirect
I’m not sold on 307/308; I think it’s probably encouraging the wrong thing, and that you want non-safe submissions to the wrong URL to fail—in fact, I’d go so far as to say that it’s preferable for these redirects to only be done for GET and HEAD requests, and return 405 on POST. These redirects are for humans that have mistyped, copied a slightly incomplete URL, or where a plain-text link detector has misperformed—all situations where HEAD and GET are the only possibilities; these redirects are not for machines except in those situations, and not for POST. POST targets are usually hard-coded or generated, so if someone gets it wrong, they can fix it because it’s obviously not working. But if you use 307/308, then you’re just slowing things down for every single user by adding a pointless redirect. (In fact, I mildly wonder if doing slash redirects is a bad idea even on GET, since the wrong URLs invariably end up in the source of web pages, slowing down the link rather than just having it obviously broken and must-fix. I think the casual cases outweigh that very mild harm, but I’m not certainly decided.) No, better not to encourage an incorrect target for any non-safe methods.
Another thing to think about here is that this kind of flexibility has a habit of leading to a lack of robustness, especially of security (Postel was wrong: <https://en.wikipedia.org/wiki/Robustness_principle#Criticism>). I can imagine a situation (inordinately convoluted and exceedingly improbable, but realistic) where the use of 307/308 allows you to smuggle something malicious in.
/path to /path/ *if and only if /path is a directory in the filesystem*
/path/ to /path/index.htm[l] (usually the default of some indexing module, often configurable so you or some other module can add index.php etc.)
Redirects can be internal or external. Nginx, for instance, does step 1 via an external redirect and step 2 via an internal redirect, so the final url displayed in the browser when the input was /path would end up as /path/ but not /path/index.html even if that's what's being served. You can, however, combine both steps into one and make it an internal redirect by doing something like "try_files $uri/index.html ..."
It's not standard to rewrite:
/path to /path.html
But almost any webserver can be configured to do so, and various websites and web apps may have reasons for doing that. Then it's merely a matter of understanding the webserver's internal configuration rule order to determine whether it'll redirect /path preferentially to /path/index.html or to /path.html
There's no universal right or wrong. Every one of those paths is different, and webservers can choose which to rewrite to which other.
Or you could just only serve it from one of the two URLs and let the other 404. (Yeah, static file servers don’t tend to be fond of working this way, but it’s not an unreasonable option if you’re forming a matrix of fundamental possibilities. And indeed, in the analysis, it’s what half the servers do at least some of the time; and it’s what all standard static file serving software that I know does out of the box.)
> Vercel, Render, and Azure Static Web Apps: slashless /resource returns content [from resource/index.html] but without redirects, resulting in multiple endpoints for the same content.
This is obviously Wrong with a capital W because of how it breaks relative URLs in the file. Surely it should just be considered a bug and fixed (probably to redirect to /resource/)?
> Almost everyone agrees that /resource should return content from resource.html
This is a biased view because you’re only considering mildly-opinionated static file servers and their configuration. If you’re serving it yourself with things like nginx or Apache httpd, you won’t get this out of the box and must opt into it. (No idea about others like Caddy.)
> (When both resource.html and resource/index.html are present and /resource/ is requested.) Netlify redirects to /resource instead.
I say of this too that it is obviously a bug. Netlify isn’t just taking an opinionated stance, it’s doing what is fairly unequivocally the Wrong Thing.
> Or you could just only serve it from one of the two URLs and let the other 404.
That's not a good idea. The URL you are choosing to remove may have built up some link equity, which you will lose if you 404 it. Instead of 404ing it, 301 redirect to the surviving URL. This allows search engines to consolidate equity from both URLs into one.
As a rule if a page has traffic it should be redirected to the next best when being removed, even if its a old promo page that's not relevant etc.
No you don't.
<link rel='canonical' href='whatever form you want your URL to have' />
For anyone wondering, in Apache this can be done with MultiViews [0], which is kind of nifty, but also a can of worms.
I find it interesting that TFA says "when resource.html and resource/index.html both exist ... Everyone agrees that /resource should return content from resource.html" but it seems like [1] MultiViews works the opposite way: "if the server receives a request for /some/dir/foo, if /some/dir has MultiViews enabled, and /some/dir/foo does not exist, then the server reads the directory looking for files named foo.*" [0].
[0] https://httpd.apache.org/docs/2.4/content-negotiation.html#m...
[1] I could be misinterpreting the doc. I haven't tested the actual behavior of httpd to know exactly what happens in a no-trailing-slash request of this nature.
But URLs don't need to correlate with filesystems -- see "clean URLs" in various php-based routers (all non-asset URLs are rewritten to load a single script). From this perspective, I agree with you.
This analogy would be an argument for allowing a trailing slash and redirecting to sans-slash, but not for requiring a slash.
For example, as far as I can know, your last sentence might have been intended to be: “In fact, when people are freed from convention in contexts like chat programs, they often drop unneeded bits like trailing periods and any semblance of grammatical decency.”
Why does an author publish an unfinished thought? Do you regularily chat with people who send in the middle of their sentences?
I absolutely fucking hate
chatting with people who
do this sort of
stupid thing.
In Ancient Greek, Latin, Hebrew, among other languages, people didn’t end sentences with periods. They didn’t even put spaces in between words. You can tell where one word ends and the next begins because you know the words of the language-and, in most languages words contain certain patterns; for example, in English, certain letters almost never occur at the end of a word, which can help you guess word boundaries even if you don’t know a particular word. Likewise, ending of sentences can be inferred from grammatical structures - although, ancient Latin style tends to involve massive run-on sentences, in part because when the end of sentences are not marked, people care less about how long sentences are. Alphabets were around for many centuries before whitespace and punctuation were invented and standardised.
if you care to browse my comments it looks nothing like this, at all, catch me on twitter dot com to see this all day erry day
I do know the local social norms but there is more than one way to illustrate them ¯\_(ツ)_/¯
https://www.nytimes.com/2021/06/29/crosswords/texting-punctu...
<link rel="canonical" href="https://example.com/resource/" />
https://developers.google.com/search/docs/advanced/crawling/...
Do you never use relative URLs?
So... I guess no more trailing slashes for me.
Some web-servers do this and others require fiddling with to get them to behave that way. But ultimately. the slash is just a separator and not semantically relevant. It's like many languages now allowing trailing commas in lists. It's convenient. The comma does not add an extra element.
One place where this comes up in practice is with specifying base URIs in e.g. configuration files. Somewhere else, this base URI is consumed to construct a full URI using a suffix that may or may not have a leading slash.
If your base URL is http://foo(/) and you want to append (/)bar, you might end up with http://foobar, http://foo//bar, http://foo/bar depending on what people do on both sides. Or those three with a trailing slash. There is no right answer.
The only sane behavior that follows the principle of the least amount of surprise is to make sure that base uri and suffix will be separated by exactly one slash and assume nothing about the presence of leading or trailing slashes. That way, nothing will break or behave unexpectedly if people add or omit a trailing or leading slash.
(Apache supports this using “MultiviewsMatch”.)
In this case it is in RFC 2396 "Uniform Resource Identifiers (URI): Generic Syntax" [1].
In Section 3 you find this. The forward stroke ("slash") is a separator:
URI that are hierarchical in nature use the slash "/" character for separating hierarchical components.
[1] http://www.ietf.org/rfc/rfc2396.txt(Your answer doesn’t enlighten anything. You could possibly be saying “it says use slash as a separator, not a terminator”, but it can be quite reasonable to declare a null component at the end as a means of accessing the default resource for that “folder”. Pages often have other resources associated with them. The fact of the matter is that the folder/file divide just isn’t present in URLs, and so the answer in a situation like this is ambiguous.)
Therefore, in relation to the article, the URI RFC does indeed not offer guidance about what would be the better approach.
Probably not.
I would say from SEO and UX perspectives this is a must.