http://news.ycombinator.com/item?id=213784
(Ordinarily I wouldn't complain, but it's happened about six times in the last 2-3 days.)
http://news.ycombinator.com/item?id=213784
(Ordinarily I wouldn't complain, but it's happened about six times in the last 2-3 days.)
http://groups.google.com/group/google-appengine/browse_threa...
and
http://groups.google.com/group/google-appengine/browse_threa...
aren't dupes, even though they link to the exact same page. There really isn't anything you can do about this particular case except error-prone heuristics.
Perhaps the "solution" is to ignore query strings....how many sites use them to distinguish content anymore? Alternately...compare the content of the <head> tag on the linked page? That wouldn't be a perfect solution, but it would probably go a long way.
Do you have JavaScript disabled by any chance, or do you use some obscure browser?
In any case, I think HN should strip the hash and what follows for purposes of dupe detection, but keep them in the link in case someone actually wants to link to a specific spot in the page.
To answer your question: query strings ("?foo=bar&a=b&c") are widely used. Among other places, HN itself uses them. :) Also, whenever you submit a form with GET.
I had forgotten that HN was using query strings to reference articles...D'oh. By now, I figured that everyone had adopted the URL-mapping approach. Anyway, detecting collisions based on the head tag still seems like a possibility....