Testing 3 million hyperlinks, lessons learned
samsaffron.com
samsaffron.com
He points to Stack Overflow's 404 as a good example and claims "We do our best to explain it was removed, why it was removed and where you could possibly find it."
Yet there is still no permanent archive of deleted Stack Overflow content; you have to rely on third party archives like archive.org and even then, you have to be lucky.
SO moderators have a habit of retrospectively deleting old content that is off-topic under current rules, even if it was perfectly on-topic at some time in the past. I feel this is bad internet citizenship -- it's removing internet history for no good reason.
Fair enough, delete newly created off-topic questions under the current moderation rules. But when these types of questions were asked originally they were on topic at the time. Deleting them retrospectively (completely - no redirect either) is still poor form.
(Top read otherwise though!)
Exclude the site and its posts/activity from most of the listings and indexes on the rest of the network, to emphasize that this content has been removed from the Stack Exchange network. Include a disclaimer in the header of every page.
Specific to this point, a new project I'm building supports "pretty" URLs and I've found my (now) favourite solution is to build an aliases system.
It works like so: when a user creates an item an "alias" is registered, it's set to "current" and all future queries to that alias are logged. If the user causes a change to the URL in future (name change, etc.) then the new alias is registered but the old one is retained and 301s to the new alias. All aliases are accessible by the user and they can invalidate them manually (if they want to re-use an alias for example) however if an alias has had a large amount of hits from a single source since that alias was retired (say 50 referrals from website.com to mysite.com/previous-alias) the system assumes that the user posted the link on another website and so invalidating that alias will cause a dead link (and lose my site traffic) so it doesn't allow it.
I guess it's convoluted and adds extra overhead but I feel like if you have pretty URLs (which are in my opinion something that a website should aim for) you need to be in a position where they're not going to cause the site to break the rest of the internet. The easy solution is to have pseudo pretty URLs (eg: website.com/123-pretty-url, where 123 = ID and pretty-url is just an ignored string) or just not allow URLs to ever be changed, but I don't like either.
I wonder if any other websites have a good approach to this.
http://stackoverflow.com/questions/427102
http://stackoverflow.com/questions/427102/what-is-a-slug
etc, will all redirect to the canonical:
http://stackoverflow.com/questions/427102/what-is-a-slug-in-...
If the title changes we update the slug and redirect with a 301 to the new canonical
Another potential solution & my preferred method is whenever a change is made that would affect the url of a page. Update a "legacy" table with the old url and the location of the new url, next time a 404 is going to be thrown do a search against the database & redirect accordingly if a new url is found. I rolled this approach into https://github.com/leonsmith/django-legacy-url and whilst it's not polished it's by far the easiest & probably most automatic/maintainable solution I have found.
Not if you properly generate and apply canonical links :)
The internet is self destructing paper. A place where anything written is soon destroyed by rapacious competition and the only preservation is to forever copy writing from sheet to sheet faster than they can burn.
If it's worth writing, it's worth keeping. If it can be kept, it might be worth writing. Would your store your brain in a startup company's vat? If you store your writing on a 3rd party site like blogger, livejournal or even on your own site, but in the complex format used by blog/wiki software de jour you will lose it forever as soon as hypersonic wings of internet labor flows direct people's energies elsewhere. For most information published on the internet, perhaps that is not a moment to soon, but how can the muse of originality soar when immolating transience brushes every feather?
Exactly. I don't understand why almost all blogging and CMS platforms store data only to database. I think that proper solution for most sites would be to keep DB for maintenance and indexing purposes. For visitors everything would be served from static files.
edit: found this: http://www.chronicleoflife.com/ ... but i was thinking something to publish instead of simple backup
edit2: probably the only company i could trust to pull this off (one-time fee for publishing static content) would be amazon. it fits really well with their core business (infrastructure), and amazon is very good with long term stuff
Currently, I back up URLs I care about or link on my site to ~3 places: the Internet Archive, WebCite, and my hard drive ( http://www.gwern.net/Archiving%20URLs )
http://theopenphotoproject.org/
Others are Unhosted (http://unhosted.org/) and OwnCloud (http://owncloud.org/).
(Might take payments off the table as well, or at least have 10-year plans or something :-))
For me one of the worst offender in this category is youtube. I can't understand why they don't put a slug with the video name in the canonical URL (especially since they have youtu.be for shortening URLs). It's really a pain to find back an old video in, say, an IRC log with only the opaque video ID.
Vimeo does the same thing. Dailymotion however does put a meaningful slug.
This story for example - http://news.ycombinator.com/item?id=4077891.
(note: this page's URL didn't change since at least 1999)
Shame as the rest of the article is quite good, but that really flags me that this is a little bit cowboy code.
Also interesting to read some sites are taking a 'white-list' approach to robots.txt, as he says this is resulting in people starting to ignore it.
Glad you liked the rest of the article, I hope this helps others
The point of blocking a link with robots.txt is to say "Hey, web crawlers, please don't load and index this page". it does not mean "Hey, users, please don't come and load and read this page".
So the script written, for all intents and purposes, is just the same as a regular old user clicking the link and reading the page then keeping a list of the links that work and those that don't. It's not a crawler, it's an automated user.
If you are a webmaster than wants to block people from posting links to your page all around the web allowing others to come and read it, make the page 403.
On the subject of GitHub's robots.txt[0], would anyone have a guess at why this particular repo[1] is singled out?
Seems unavoidable on large sites.
1. sys-admin reorganisation, moving content from one spot to another without redirects in place.
2. developer reorganisation, for example moving from "confusing" urls to "slug" based urls without adding redirects
3. fragile content, content that moves depending on external changes (beta to release for example)
4. product retiring or companies getting acquired
5. hackers messing stuff up in a way that can not be fully repaired (or an non-recoverable data loss)
2. list mirrors like nabble do wholesale migrations without redirects (Google groups is gong thru this now with new format (but with redirects, I hope):
groups.google.com/forum/#!msg
3. wiki's get pages duped and branched to where their marginal utility is 0, so the sponsor decides to start over.Often it's clear that what's there now is a lot better or more professional, so you can see why the person didn't feel like messing up the site with the old stuff - but that old stuff is still gone.
Other times the domain is gone, since the person or people have moved on to doing a lot better stuff and stopped maintaining that old site - "why bother." People change - a web site isn't something you publish once, it's something you publish every time your server answers an http request. Would you keep publishing everything you wrote 10 years ago?
how about "fuck you"? I guess it's high time to make honeypots, tarpits and bans common practice.
No, we are not going to ban all the links from GitHub on our site cause shitty WEB CRAWLERS forced GitHub to use a white-list based approach. This is not WEB CRAWLING. It is link validating. We are not crawling in the sense of building a huge tree of links. We are testing that the external links on our sites work. If we are not allowed to test them, why are our users allowed to click them? Are we not committing an even greater crime by allowing these links on our web site?
"The Robot Exclusion Standard, also known as the Robots Exclusion Protocol or robots.txt protocol, is a convention to prevent cooperating web crawlers and other web robots from accessing all or part of a website which is otherwise publicly viewable."
The convention is a best effort thing, we tried to respect it, but doing so was both AGAINST what the authors of the Robots.txt file at GitHub intended AND the spec is advisory, not an IETF RFC. If it was an RFC then some smart people would review it and turn it into something sane and usable that deals with this exact use case.
And you know what, an RFC would NEVER pass for robots.txt as it is now cause the white-listing potential is anti Internet. Why should Google and Bing be the only parties who are allowed to discover content on the Internet? User agent restrictions are completely evil, wrong and backwards.
Sorry to shatter your imaginary delusion of what you think the Internet is.
Followed by shitty" and "web crawlers" in all caps ^^ Someone ate a clown for breakfast I see.
"No, we are not going to ban all the links from GitHub"
What? Slow down there -- why would you care about invalid links? Did you just say that you can't possibly allow users to post links, as long as you don't know they work for automatic crawlers, not just for human visitors? And someone else chimed in saying you give arguments? Heh.
Well, you give an attempt of one, with "we are not crawling in the sense of", and then refute it with the bit you quoted: "web crawlers and other web robots". It's not called webscraper.txt, it's robots.txt period.
So how then would a website determine a rogue user agent? You dress up like the slimy guys, you get the banhammer -- what do you expect? If you care so much about Facebook and Twitter "content" that it is worth it for you to be undistuingishable from attackers, then just cope with it. But don't pout at me, just eat up what you ordered.
And what delusions about the internet? You just beat around the bush and then finish with that strawman? And what is an "imaginary delusion", by the way? The one you imagine I have? Now that's a Freudian slip if I ever saw one ^^
"Completely evil, wrong and backwards"... so... You're entitled to know the validity of links posted on your site, but website owners aren't allowed to care about their resources and who they offer them to? Who's deluded?
Because Stack Overflow is a site whose purpose is to answer questions. People may provide links when asking or answering a question, and those links may be important in understanding either the question or the answer. So invalid links degrade the value of the site.
What they're doing is fundamentally different than web crawling. Web crawlers are about discovering content. That means starting at a root and crawling out to see what you can find. One URL can spawn many more URLs to look at. They are starting with a known URL, and seeing if they can visit that URL. They have one URL, and only visit one URL.
The problem with whitelist-only robots.txt is that they favor monopolies and startups are the ones getting the "fuck you". But maybe you don't care about that.
These are supposedly good guys. So my reaction was "You gotta be fucking kidding?! You didn't just say that it's inconvient how some sites use robots.txt, so you just throw it out altogether for your precious little bot and epically important link checking quest. No wait, you did. Oh well then, BYE."
Oh well. I guess this is hack news, not hacker news, my bad :P
User-agent: *
Disallow: /
in it?
What's the difference?