Cool URIs can be ugly (2023)
unterwaditzer.net
unterwaditzer.net
That moniker, at least for me, is reserved to links that look something like this: `/{endpoint}/{long_hash}?__gtr[0]&__jd__[df]=%ezaz54%d/{another_very_long_hash}[c__f]/`
Now, that's an ugly URL!
It's hard to find an example Of one of those now, because the kind of sites that tolerate weird comma-infested URLs in 2003 aren't the kind of sites that meticulously maintain those URLs in working order for 20+ years!
Right now, I’m knee-deep in coding my Django app. I totally dig how the framework kinda "forces" you to write neat URLs ― it’s one of my favorite things about it. This might seem silly, but I actually take immense pride in crafting simple, elegant URLs, even if the majority of the users won't even notice it.
As for the comma infested URLs, the website of one of the major news outlets in my country manifests such behavior. It always puzzled me as to what tech stack they were using. I'm not sayin they still use it today (as Vignette went belly up in 2009), but this can be a heritage from those days.
I really enjoy using Django since I first got to know it back in the 2.2 days, I’ve used nothing else for my projects, big or small. I’m head over heels for every bit of it and having recommending it for years to my friends!
Big thanks to you, Simon, for helping create this awesome piece of tech!
Example of a "correct" url
?value=A&value=B&value=C
Complete frameworks would have a method that returned the values as a list. Some like PHP required ugly work arounds where you had to name the parameter using the array syntax: value[]=A&value[]=B&value[]=C
Even if the framework supported multi-values, many preferred the shorter version: value=A,B,C and split the values in code instead
values = request.GET.getlist("value")
# values is now ["A", "B"]
We built it that way because we had seen the weird bugs that cropped up with the PHP solution, where passing ?q[]=x to a PHP application that expected ?q=x could result in an array passed to code that expected a string.Handy if anything actually supported it, because then you could plop parameters on the end of urls without looking to see if there was already a ?
I'm sure it's used somewhere, but I can only remember it being used by Yahoo Link Tracking ;_ylt=y64encodedgunk
So, people who learnt programming in 2000, until ~2010 it's quite normal to see the commas as delimiter of multiple parameters.
The only Cool URI failure on Wikipedia is the ".m" which is added to the mobile view.
https://hudoc.echr.coe.int/#{%22documentcollectionid2%22:[%22GRANDCHAMBER%22,%22CHAMBER%22],%22itemid%22:[%22001-230857%22]}
Fortunately, most pages list a "clean" URL that also works: https://hudoc.echr.coe.int/?i=001-230857It’s not much harder to hide it; E.g. for static files, create a directory and put an index.html there.
That the page is sent to your browser as HTML is not a defining attribute and could very well depend on HTTP content negotiation.
Yes, why not? Just because file extensions matter to certain systems doesn't mean they do for others, and nothing about a URL to a file is required to match its DOS/Windows friendly file name.
> GET /<artistname>/<albumname>/<songname>/download HTTP/1.1
> Host: fakemusicstore.example
< HTTP/1.1 200 OK
< Content-Type: audio/mpeg
< Content-Disposition: attachment; filename="<artistname> - <songname>.mp3"
TBH I agree, I personally do my best to ensure the extension in the URL matches the document type on sites I run, but my point was that it's not in any way required and it's actually somewhat common for it to not be the case where the person I was replying to seemed to think it mattered.
The hiding of the index file name in a folder is kind of a quirk, it's automatic behavior that is being taken advantage of to make "nice" URIs, but it's actually hiding useful information.
File extensions, while a DOS/Windows thing, I've found to be an extremely useful convention on unix, and linux, and just about any other system I've used (though I can't remember what we did on VAX box we used to use in the 90s).
https://m.facebook.com/story.php/?id={redacted}&story_fbid={redacted}The popular "Cool URIs don't change" article at <https://www.w3.org/Provider/Style/URI> says:
> What to leave out
> ...
> File name extension. This is a very common one. "cgi", even ".html" is something which will change. You may not be using HTML for that page in 20 years time, but you might want today's links to it to still be valid.
But I have been running my website for over 20 years now and I do think I'll stick with ".html" for the foreseeable future. This combined with the fact that I strictly use relative links for cross-linking between pages, for loading CSS, images, favicons, etc. means that I can browse my website offline (directly from my local disk) too just by opening the local index.html file on my web browser.
(Apart from supporting optional extensions, this code also supports throwing an error if someone prepends dots into the url - which, for me, indicates someone probing the server for weaknesses and is not a legit request.)
The funny thing is that I still often use file extensions since IntelliJ can only let me easily navigate/check existence if I use the extension.
Eventually I'll support slugs in the filename by just ignoring everything after the first dash.
Personally I'm a fan of including a post ID in the URL, e.g. /category/123/post-name. Because if you want/need to change the URL later, you can simply parse the URL to get the ID back to create redirects. A lot of sites of all scales don't implement redirects which makes me sad.
I think there was a news site acquired by Bloomberg, I forgot the name. When you visited an article in the old domain, it redirected to a landing page on Bloomberg saying it was part of Bloomberg now instead of redirecting to its new URL.
You can thank the browser complexity moat for that. If browsers were simpler to implement someone would have started experimenting with this (markdown at least) years ago and other browsers would have picked it up.
Asking genuinely, I don't know, but it's an important fact to take into account if you're planning ahead.
Using plugins, you could think about Markdown, wiki markup, ...
Similar type handler is engaged with XML. Unless you can utilize W3C standards to implement a custom markup language using XML/XSLT and have it work across browsers without plugins.
SVG is vector graphics.
For another full markup to be even considered there would have to be one that's widely adopted and realized through plugins. Nobody is making interventions in standards to open up venues for easy implementation of custom markups when those markups are used by 0.001% of publishers.
I don't see why that would change.
If browsers were easier to make, someone could experiment with content negotiating for markdown and rendering it client side.
The stuff you're talking about isn't about browsers its about the websites.
If you had a website that uses javascript to parse MD or any other markup, spit it out as trivial HTML with light DOM, client-side formatting can do everything you want.
The problem is that modern websites use patterns that workaround users' capability to customize the presentation of the website. They do not want you to look at their site the way you want.
I guess the best solution to that isn't browser-side .md rendering, though.
I don't expect that url.html is a static html file. I expect it to be server-side generated in 2024. For me site.com/page and site.com/page.html are the same. I do not expect different behavior from my web client side. So I may switch backend engine every year, and I'll just route the request sfrom page.html and that's it.
What's way worse than this is using non-HTML extensions for emitting html. I go to pichost.com/image.jpg and I get a webpage served. This is a bad pattern and it needs to go away. I'm not even going into responding differently depending on user-agent or referrer, if you have combination of these you get JPG returned, if you don't you get a webpage returned.
It's mostly based on the Accept header these days (browsers don't tend to include HTML there in image contexts) and the Referer should have been removed decades ago. This means browsers (the ones with a large market share at least) are 100% complicit in enabling this behavior.
HTTP has no concept of a file extension.
Needless complexity if all you need is best served by a static html file.
There is an escape mechanism for making endpoints without an extension but we rarely use it.
It’s a weird design I probably wouldn’t make these days, but for debugging at a glance it’s honestly pretty nice to look at the stream of requests and just know the type of each.
There's no excuse for not implementing them properly, however. I'm less of a fan of the existence of verbs, which I consider to be a part of the URI which isn't in the URI itself. Things would be better if one URI was one endpoint, rather than potentially as many endpoints as there are verbs.
That's /blog/slug/ which should return the default file for that directory or generate an index (ahem) of what's in that directory.
./slug.html <-> /blog/slug
./slug/index.html <-> /blog/slug/
> Pages will also redirect HTML pages to their extension-less counterparts: for instance, /contact.html will be redirected to /contact, and /about/index.html will be redirected to /about/.
https://developers.cloudflare.com/pages/configuration/servin...
Keep in mind that this is Cloudflare Pages, not Cloudflare in general. Cloudflare Pages is a product where you give it a bunch of files, and it serves them as a web site. You don't have your own server behind Cloudflare in this case.
Serving a web site based on a directory of files is tricky, because URL space and filesystem space are a little bit different. Files on disk need to have file extensions to indicate their type, but URLs are not supposed to have file extensions, because their type is indicated by the `Content-Type` header. So if you are taking a bunch of files and serving them as a site, you need to figure out how to transform the type info in the URLs into Content-Type headers in an appropriate way. This is a solution to that.
Another remapping that nearly every file-based web server does is, if the URL turns out to be a directory, it returns a redirect to add `/` to the end, and then from there it serves the file called `index.html` in that directory. Again, this is needed because URL space and filesystem space don't exactly match: a directory on the filesystem cannot itself have byte content, it can only contain files. But a URL that is a directory can also directly serve content, so you have to figure out how to resolve that.
`index.html` remapping is pretty much universally accepted. But it's true that people have differing opinions on extension-stripping. The extension is redundant, but some people would rather keep it just to make it clearer how URLs map to files. Fair enough.
Unfortunately Cloudflare Pages does not have a setting for this right now. It has chosen to implement only the most popular approach. This is a product decision, and of course some people will disagree with it. You can submit a feature request, or you can use a different product that works the way you want (there are tons of them out there). But it's not a "bug" that the product has not chosen to implement your specific preferences.
(Disclosure: I work for Cloudflare, but not specifically on Pages.)
There are probably settings for this stuff.
It's quite normal for static sites to do something weird here. Like having folders that all contain index.html, and then having settings to strip (or add) the final slash.
There are so many different flavors, the only somewhat neutral default is what apache does.. still it's not much :)
Including ".html" in the URL when you're first creating a site signifies a risk that it'll change in the future, because it's evidence you went along with what was easiest to get the backend technology to serve your content, and as the backend changes over time, you'll do that again, changing the visible URI as you go and causing bitrot.
But if you picked ".html" and stuck with it, that's now the cool URL, and you should use web server configuration to make sure it remains that way, even if the backend technology has changed completely.
This is how I decided to configure my nginx as well for my web page, but note that it still locks you into something: you will still end up seeing links out there that reference /path without the extension and you will need to set up all future web servers to find the right resource on that URL. (Even if that is by adding files to the file system rather than writing web server configuration.)
Here's pong and epic of Gilgamesh in a Tweet:
https://twitter.com/rafalpast/status/1316836397903474688
Context: https://sonnet.io/projects#:~:text=Laconic!%20(a%20Twitter%2...
Look for "Laconic! (a Twitter CDN)" in the project section.
(Also, weirdly enough, fragment links can be consumed in Safari, but to share them, I need to open the site in Chrome.)
[1] https://www.rfc-editor.org/rfc/rfc2396
[2] https://www.rfc-editor.org/rfc/rfc1737
[3] https://www.rfc-editor.org/rfc/rfc1737
[4] https://en.wikipedia.org/wiki/Uniform_Resource_Characteristi...
[5] https://en.wikipedia.org/wiki/Uniform_Resource_Name#URIs,_UR...
Or maybe the presence of host:port qualifies it as a URL.
One type of URI is a URN (name), e.g. doi:10.5281/ZENODO.31780 - a unique name for a resource, but no instructions on how to obtain it
Another type of URI is a URL (location), e.g. https://doi.org/10.5281/ZENODO.31780 - same resource in this case, but now we know we can obtain it via the HTTPS protocol
Few people call the address in the web browser a "URI" any more, even though technically it is one. Your JDBC URL is a URL, as is "mailto:president@whitehouse.gov" or "tel:+44-118-999-881-999-119-7253"
Yup, URNs were part of the "semantic web" craze, so you could e.g. record facts about a book with isbn: scheme URNs. Nothing much consequential ever came from all that committee busywork, but people got to pontificate and sound smart talking about reification and so on. I still wonder who paid for all of it.
https://en.wikipedia.org/wiki/Resource_Description_Framework
A URL is a URI.
I know it's fashionable to have URLs like:
https://example.com/appService1/getOrder/17
But it's hard for monitoring tool to tell, which part of URL is API endpoint (which you want to report on) and which is user data (which you don't want to report on). I wished people used query portion of the URL for user data, so it's syntactically distinct from the path.
github's network effects are far more insidious.
I disagree - this still requires you to support both in future unless you are happy breaking old links.
In general trying random URLs and them accidentally working and then not working, despite you weren't linked from somewhere is not something that counts as a broken link.
Say for example you added "?page=123" to a URL that had no pagination. So the normal page opens but it ignores the parameter. Then later the parameter is added, so when you add this parameter now you get a 404, because there's no such page. Was a URL "broken"? No.
Erm I mean "Cow".
The TLDR version:
/foo will serve content from foo.html, if it exists
/folder will redirect to /folder/
/folder/ will serve folder/index.html
A 404.html file will be used for 404s
The .html rule beats the folder redirect rule
index.json works as an index document too
If there is no index.html or index.json a folder will 404
> But Cloudflare’s redirect is permanent and has been public for a few weeks, therefore all Google search results were pointing to the cleaned up URLs. If I wanted to move to a different static site host, I would have to install additional redirects so that none of those links break, just to clean up a mess I didn’t cause.
The "would have to" remark is odd. It's too late; you'll need to install redirects to stop those links from breaking anyway. Whether GitHub supports this automatically doesn't change anything. You may as well have not switched.
Many static content hosting services have this exact behavior. In fact, many web servers have offered this behavior, going back decades, because it's what a lot of people want. It's kind of needed to work around the fact that files usually indicate their type by filename extension, but URLs are not supposed to have such extensions since they indicate their file type by `Content-Type` header.
(I work for Cloudflare but not on Pages specifically.)
No, just because I hosted something for awhile does not mean I am obligated to host that resource in the exact same way for eternity. There is no contract, implicit, social, or otherwise that I will continue to provide that free thing for you in a way that is convenient to you personally in perpetuity.
Oh, you didn't perfectly lay out your URIs in the initial design? Too bad, you're saddled with the unending burden of maintaining redirects forever or you're not "cool". Should have known the company was going to move to Markdown static site generation five years before Markdown was invented.
Miss me with that shit. Link rot is the burden of the link author, not the target.
<meta http-equiv="refresh" …>
sent in the head, and some html/css to make it pretty. It's not ideal, but I assume search engines support it (dunno if there's any additional SEO improvements).Right because keeping a list of source->destination and configuring your current server based on that is such a burden...
> Miss me with that shit. Link rot is the burden of the link author, not the target.
The link author isn't the one making the changes, the target is. The link author might not even be alive anymore. Expecting others to untangle your mess is ... not cool.
Okay, but you did know, right? Maybe not that the new thing would be called Markdown or exactly when but that there would be a new thing. The W3C sure knew and told you. That's why they wrote e.g. this paragraph:
> Software mechanisms. Look for "cgi", "exec" and other give-away "look what software we are using" bits in URIs. Anyone want to commit to using perl cgi scripts all their lives? Nope? Cut out the .pl. Read the server manual on how to do it.
> Pretty much the only good reason for a document to disappear from the Web is that the company which owned the domain name went out of business or can no longer afford to keep the server running.
speaking of... wonder if there are any 302 redirect managing software...
Mind you, this feature is still under development, but this is the ultimate goal of my app.
It is currently in free beta if you are interested in giving it a go: https://bernard.app
I mean sure but… be cooler if you did