GitHub: a case study in link maintenance and 404 pages
chrismorgan.info
chrismorgan.info
All we need is an error code 402.5 "plausible deniability between unauthorized and not found"...
For the rest of your users it wouldn't hurt to say, "You might see something different if you're logged in." Or if they are logged in, saying, "Were you expecting something different? Maybe you just don't have access yet."
You can press 'y' to expand the URL to its canonical form.
I don't think the answer is forcing the canonical URL on every page request, though. I think it's important to retain branch names instead of the full sha. I find this, for example:
https://github.com/holman/dotfiles/blob/master/osx/set-defaults.sh
…far more usable and meaningful than this: https://github.com/holman/dotfiles/blob/0fe9e9963b2389eae4c9de49a4873bd819e19067/osx/set-defaults.sh
It'd be great to be able to support file renames and deleted branches better in the product, but that takes some time to build out. Hopefully we can do better with that in the future (and we have been working on things like this recently- supporting repository redirects was a huge one, really).https://github.com/takezoe/gitbucket/edit/master/README.md
you will see a 404 if you're not logged in with no information.
At least I gave http://relink.chrismorgan.info a proper 404 page. The sentiments displayed on it are the same as what I would put for http://chrismorgan.info, though: you won't get a 404 page on my site unless you broke the link yourself.
This sounds all well and good in theory, but - commercially speaking - becomes very impractical for many businesses. I'd be really curious what the ROI would be on a more rigorous link maintenance practice. My general impression is that it's simply not worth it.
> I should switch it to Apache so that I can get a 404 page, but I haven't got round to doing that.[1]
Interesting comment from the OP, whom I suspect also recognizes (quite possibly like GitHub) that the time it takes to do this type of stuff hardly outweighs the benefit.
I also think it's easy to underestimate the R here. It's very hard to detect things like brand damage and lost leads, especially when somebody turns up on your site once, thinks you guys are chumps because of a bad first impression, and never comes back. Whereas the I is obvious, because the cost is all internal to the company.
The two factors, alas, reinforce one another. Link rot isn't an obvious problem when setting up a system, because nobody is linking in. By the time the problem becomes noticeable, the system is hard to change, making the cost of a fix high. And that trains people to treat link rot as unimportant, deepening the cycle.
I put things on the web so people can see/use them. If they come to my site, I want them to get what they're looking for. If I'm going to serve a 404, it's either because a) I screwed up, or b) somebody mistyped something. In either case, distracting them with something else can only take them farther away from their original goal.
EDIT: To be clear, this is more in regard to the original articles qualification on the permanence of links: If for some reason it can’t exist any more, don’t just let it go: it should show useful error. Using 410 GONE instead of 404 NOT FOUND after a DELETE is a fairly minimal but direct application of this principle.