Cool URIs don't change.
w3.org
w3.org
It can be much worse though...exposing machine names, unnecessary complexity and parameters, all changing from year to year.
A server doesn't have to puke its implementation details everywhere. I can't even count all the "enterprise" apps that make that mistake (e.g. a helpdesk system that gives me "serverNameThatWillChangeNextYear.domain.net/some/unnecessarily/convoluted/path.unnecessaryExtension?whatthehellisallthis&garbage1=a&garbage2=b&finallyRelevantBugNumber" instead of a stable URL like "company.net/bugs/bug123456". And just try E-mailing a complex URL to somebody (it wraps, and time is spent to awkwardly correct it).
Honestly, it's as if most web developers don't understand how powerful URLs can be. If you make URLs short and stable and use them to help look up stuff, they can be very nice. Instead in the past I've seen people E-mailing 9-step instructions on how to find something because the damn URL is unreliable.
Or, more likely, most of the "enterprise" apps with those URLs were first released over a decade ago, when the idea of "pretty URLs" was not yet mainstream. There was not even a mention of mod_rewrite on Wikipedia at that point. The article on the front controller pattern, which most webapps that handle their own URL routing use, wasn't written until 2008.
Or, even more likely, the enterprise customer never specified "URLs that aren't terrible" in their 100-page RFP, and so the lowest bidder didn't bother with URLs that aren't terrible.
HTML, and server paths / CGIs were largely hand-coded and short.
The URI explosion occurred in the late 1990s / early aughts for the most part with Java and Microsoft entering the fray from my recollection.
Let's say I have an old-school, all-static site with pages at http://example.com/x/index.html and http://example.com/x/about.html. I would like to make a link to the "index" page from the "about" page. What are my choices?
<a href="/x/"> will work, but will break when someone decides to move "/x" to "/y".
<a href="."> will work from the server, which does an internal redirect, but not on the static version on my local drive I'm going to demo to my boss. (I also have a hunch a significant number of developers aren't aware of "." and don't know this is an option.)
So we end up with <a href="index.html"> for better or worse.
<a href="/x/"> will work, but will break when someone decides to move "/x" to "/y".
If /x redirected to /y, that wouldn't be a problem.99% of the time the reason for index.html is the developer was viewing static files in their browser because they didn't have a proper server or test environment. This is inexcusable in 2012.
Eg: http://www.example.com/path/to/url/ should open to the first specified of the the specified default objects, typically index.html index.shtml index.php index.php, or similar. This is defined in your Apache conf file, or locally via .htaccess.
If you want to refer to a specific non-index page, you'd specify http://www.example.com/path/to/uri/somepage.html
If I change my file formats from HTML to SHTML, .jsp, .php, etc., the URL changes.
If I change from files to directories for every possible URL instance, the URL changes.
The user shouldn't care.
So, is this presenting an implementation detail as well?
I only asked my original question because I thought I was missing something...
If you'd started with the 'document'-.'html' naming convention, you could use any of numerous webserver hacks to preserve this illusion. How you request something has little to do with what the server does to satisfy your request.
What I was distinguishing earlier was the distinction between 'return default index of this level of the path' and 'return a specific document from this level of the path'. Specifically indicating 'index.html' is a tad gauche.
So, regarding the "index.html" portion: The index is the default. There's no point having a default if you have to specify it, and there's no point specifying it if it's the default.
If I do find myself working with LAMP (which a lot of people still use) then I use this option, and my uris magically no longer end with ".php"
I agree with your point, but it's also valuable to understand the structural reasons that prevent such things.
OS X tries to hide extensions by default, which like the behavior in Windows that's similar, seems dangerous. Too many times people have been stung by sexy.jpg.exe.
If it's bad magic, update your distro. If it's a Konqueror error, file a bug. I'd be surprised if upstream hasn't addressed this (quick DDG/Google doesn't turn up any similar complaints).
Just like .css and .log are text files, but I may not want the same default file handler.
What I like is having *.epub open in, say, calibre, but if I append a .zip extension then it will open in xarchive.
Anyways, I open most things from the command line and wrote my version of 'open' so I get what i expect 99% of the time. :)
I'm using KDE3 so I'm not expecting any upstream fixes for this in my lifetime. I can live with crafting my own solutions.
What might work best is if mime types were used by default but forcing behavior for specific extensions was much easier. Get the best of both worlds (which I can mostly do in KDE3 Konqueror but it's tedious.)
Other examples: tar.gz, WAR files, most ODF formats.
Edit: smartphone tyops fixed.
Debian added upstream epub support (and backed out its own) in September, 2011, so this is a pretty recent feature.
Or do you mean storing a database of the application associated, specifically, with any given file (so that I might open, say, a given myprog.c with vi, emacs, or textmate)?
The Mac OS resource fork model attempts the second, but it's an inherently single-user concept that leaves artifacts around for other users and/or on shared media in an annoying way. A shadow filesystem maintained on a per-user basis under their control would be a preferred solution.
I remember when http://microsoft.com/ began doing external redirects to "default.asp" circa 1997. If you were a "webmaster" (do these exist anymore?), this was a dog whistle. They were not using static .html (or .htm) but not any of the common dynamic methods like .cgi or .shtml either. And using "default" rather than "index" indicated a break from NCSA/Apache convention. They were using a different web server. Those 11 extra characters said a lot.
First, holy crap, I hadn't been to http://www.w3.org/ in a long time, and it looks like they've actually made it to the 21st century!
Second, perhaps cool URIs don't change, but it seems like http://www.w3.org/Provider/Style/URI.html is kind of an unfortunate URI. What's "Provider"? Why are "Provider" and "Style" uppercase? And what's wrong with 301 redirecting (don't break old URIs, but still restructure them as your website matures and you realize a better organizational hierarchy)?
Third and perhaps most importantly, this all seems like a pretty awesome problem to have! How many websites survive more than a few years? (Geocities doesn't count.)
They're uppercase because they're uppercase. Asking why is like asking why Rubyists like_using_method_names_like_this and .NET devs LikeUsingMethodNamesLikeThis.
Oh, and I am not a Ruby developer ;-)
As for a zigzag effect I think CamelCasing is at least that bad, maybe moreso.
Wondering if I get points for bringing early Medieval history into a discussion on coding conventions ;-)
word+empty-space+word
but
word+line-down-below+word
so your eyes go zig-zag from the taller glyphs of the word to the low glyph of the underscore (vertical zig zag motion), instead of jumping to the next world (horizontal).
Searching for the source now.
It's because it is a part of a collection of numerous pages named "Putting Information onto the Web". Its audience is obviously (content) providers, which the URI reflects.
Take a look at w3.org/Provider and w3.org/Provider/Style. It's actually a well-constructed URL. It appears to be a "Style guide" for "web content Providers."
http://www.w3.org/Provider/Style/URI
works and is more in the spirit of the article than
> [...]
> Consider linking to your home page and/or site map [...]
For a second I thought you were advocating redirecting 404s to the homepage, and I was about to flame you. I hate that one, and I think it's harmful enough that you probably ought to have a point about not doing it in that article.
Top 10 mistakes of Web design http://www.useit.com/alertbox/9605.html
More classics anybody?
Downvotes are for comments that don't contribute, not comments you disagree with.
The workaround is the Wayback Machine, which is amazing, but could be more comprehensive and frequently updated. I wish someone like Google would throw more servers at it.
All that said, you're fundamentally right -- sometimes information stops being available because it's out of date, and keeping it available would be confusing (if a product is no longer available, it would be strange to maintain a page describing it for years afterwards). Archiving through the Wayback machine is a very helpful stopgap, but expecting them to continuously archive every version of the entire Internet for all time won't scale.
What's needed is a distributed, decentralized system, ideally at the protocol level. Imagine if a GET request by default gave you the "current" version of a page, but you could send an extra header that said "give me this page, as it appeared at date-time X". This would remove the confusion caused by the existence of a page being conflated with that page being current[1], and allow sites to maintain clean navigational and data structures by flagging outdated pages as "expired" instead of completely deleting them. When a server got a request for a page that used to but no longer exists, it could respond with a new 4xx-series header, "No longer current", indicating the document is not available for the given date-time, but is available for an earlier date.
[1] I frequently get people sending me ANGRY emails about flippant, immature blog posts I wrote 10+ years ago[2]. They assume that because it's still on my website, I still stand by those statements, when in fact I'm just reluctant to delete information.
[2] The posts still get traffic, because links to them made 10+ years ago still work, despite rewriting my CMS 3 times.
Sounds like Freenet USK's (see https://freenetproject.org/understand.html, search for USK (and boo on them for not having any anchors on that page)).
It's your lucky day: http://www.mementoweb.org/
Couldn't you implement a clumsy manual version of the 'expired' header you're proposing, by doing something like having your server precede each page with "The following has not been modified since …, and should not be regarded as current" if it is more than a certain amount of time old?
For example, is there any harm in changing where that contact us form points to? Do you have to maintain the contact us forms API to be backwards compatible forever? This really depends on what you are doing.
I agree with the article as it relates to documents for the most part. There are cases where maintaining URL backwards compatibility is a problem due to unforeseen dead ends. However, for the most part it shouldn't be. However, URIs needs for persistence varies so widely that I don't know one can generalize much beyond that.
The point could more be made: "Why not keep that URL the same?"
It's trivial to do technically, and there's no advantage to be had from moving back and forth between foo.com/contact and foo.com/contact-us. If the URL of your company's contact form gets published in a book, would that change your mind?
But where it is a form submission target, the only reason to care is if you are accepting third party form submission. But we are no longer talking about hypertext resources at that point which is my main point.
Maintaining in perpetuity a complete set of redirects (or stack of rewrite rules) for every page for a website consisting of millions of pages for decades is not feasible in most cases. Things change, departments are created or disbanded or renamed. Management decides certain things should not be accessible to the public or stored in a particular location.
It's incredibly unrealistic to expect every URI to be permanent.
Nope, it's the lack of a crystal ball. The technology world moves fast. And breaks things, as Zuck says. I have no idea how a site of mine might be structured even a year from now, and if the currently URL scheme will conflict directly with its needs or not, or if an old-to-new translation will be trivial or overly resource-intensive. By now, the world has mostly realized that over-planning is bad, and agile is important.
So who cares if a 5-year-old URL to a page virtually nobody ever visits anymore doesn't work. It's easy to Google the keywords in the link, and you'll probably be able to find the content if it's still around.
(Obviously big sites have an incentive to keep their links working, but they don't need an article from the W3C to tell them that.)
And nice strawman, by the way. I have had some EXTREMELY popular blog posts that I bookmarked (remember NVIE’s post on git branching?) 404 just because the site changed architecture. (Now hosted on GitHub.) It’s all laziness and lack of forethought and planning. What is agile about one of the most popular blog posts in the development world 404ing?
That said, yes, it can be a bit time-consuming to map old URLs to new. I once started a project that would help considerably with this, but my employer at the time decided it wasn’t worth it (hey, would they see any more gold coins if they helped clients’ visitors’ bookmarks and search results work?). But a manual connection isn’t impossible. In some cases, it’s easy. When I moved off Drupal (damn it to hell), I redirected all those old /node/<id> URLs to their new ones. Because, damn it, I wasn’t going to add any 404s to the Internet if I could. I didn’t blog to make money. I blogged to help people. And maintaining URLs helps people.
It's the UX inertia of having to write mod_rewrite rules that ensure its never a priority and thus never happens.
That is your mistake. The URI should be decoupled from the technology you are using. The URI is part of the content and not of the software behind the site.
If your stack requires tight coupling with URL structure, you have doomed yourself already.
I love how this starts out by (wrongly) assuming that "you" is a single person or even stable group of people, and that the answer to a changing URI is to say "You did it wrong."