War on Urchin
blog.pinboard.in
blog.pinboard.in
The very long URLs created by Urchin and other web analytics was one of the problems that URL shorteners were created to solve.
Whether short URLs have a "malicious" effect is a topic for a different discussion, but whatever the effects of URL shortening, Urchin parameters are not among them.
Urchin parameters certainly do make it harder to detect duplicate URLs, which would be pinboard's primary problem with them. However, so lots of other URL parameters like sessions, landing page refs, etc.. Urchin's are just the most common ones.
[Disclosure: I work for awe.sm, which provides social media analytics using, amongst other methods, short URLs]
[Disclosure: I hope your entire product category dies]
I'm arguing that URL shorteners have had nothing to do with the popularity of Google Analytics, which is the source of urchin parameters. The dominance of Google and the massive profit incentive around accurately tracking adwords and adsense is what's driving that. Consumers don't care what their URLs look like, so it doesn't matter whether URL shorteners obfuscate them.
I recognize that shortened URLs are troublesome for a bookmarking service, but they are hardly insurmountable. I'm not sure I understand why you'd have so much hate for us.
(And for the record, URL shortening is not our primary product)
I saw URL shorteners first in a paper magazine, the articles about Internet provided the links in that format (tinyurl.com), much easier to write in a computer keyboard.
I can think of quite a few other support frameworks for content creators -- some within the capitalist system, some outside it.
(My take is that almost all advertising sucks, from a consumer point of view, because it's designed to steal the consumer's attention from the item that drew it in the first place. At its worst, it becomes as unwelcome as spam. After all, we've only got 168 hours in any given week to pay attention to stuff: ads steal the only thing from us that money can't buy, and that's time.)
Sounds reasonable to me.
The truth is that outside of nerd circles most people don't even understand URLs, far less care whether they have extra parameters in them. Major browsers are considering getting rid of the URL bar entirely. The tidiness of URLs is therefore of almost zero concern to major online publishers, but accurate analytics is, which is why UTM tracking is so popular.
These parameters have nothing to do with serving the content the user wants and are only there to track users and behavior and are metadata about the real url that's attached like a parasite. I understand the value for content providers, but I think stripping them off for archival purposes is appropriate.
In a world without analytics you have tons of crappy content written or created at the whip of an executive who thinks she's good at guessing market demand (and nobody in the company can prove her wrong scientifically speaking before the job is done). That content will prove to be a failure when it's launched in let's say 9 out of 10 cases, which means 9 bankrupt projects, 10 times less interesting content on the web at 10 times higher costs of production, which in turn leads to less competition, higher prices (paywalls anyone?), reduced rate of learning/innovation etc.
I have a balanced view on the issue and I know the pros and cons of each side, including the privacy issues involved for everyone when surfing the web. But I'm sad when I see remarks as "I hope your product dies" or when someone chooses to blatantly represent just one side of the story.
That's a completely one sided and biased portrayal. They also help webmasters manipulate the psychology of users, hinder privacy, make an open web more difficult, etc, etc.
Content providers already have feedburner for rss metrics (also, now by google) not to mention google analytics (nee urchintracker) and good ol' fashioned server access logs (which you can analyze with urchin proper (or mint, or what have you)).
I can see from a gut-reaction standpoint how you could write what you did, but aside from gawker (who notoriously uses analytics, c.f.[1]), how many legitimate content providers use analytics as anything other than a rough barometer for trends? The problems you describe are problematic for certain types of tabloid publishers (drudge report, ny post, etc.), but they are hardly addressed by a handful of querystring parameters.
[1] http://www.newyorker.com/reporting/2010/10/18/101018fa_fact_...
Nowadays it's built into Adwords, so the answer is anybody that sets up conversion tracking in his Adwords account.
More details: Google provides a tool called conversion optimizer[1]: it's enough to put a tracking code on one of your objective pages (the purchase page, the signup page etc) and Adwords will use machine learning to see the analytics for ads that convert well to your objective (what keywords did they use in Google, what locations are they coming from etc). This way, you can stop paying money for keywords with 20 clicks and 0 conversions, and instead you can raise your ad bids on those keywords performing well (i.e. 5 clicks and 3 conversions). The publisher is happy (more conversions, less clicks, less money), Google is happy (better targetting means less impressions used means more impressions remaining in the inventory to be sold to others for additional income), the customers are happy (publishers with lower customer acquisition costs can pour more money into the actual content/product).
[1] http://www.google.com/adwords/conversionoptimizer/howitworks...
The utm parameters allow a site to track campaign information and replace much more annoying techniques like setting up unique landing pages or redirects. There's no relationship between these parameters, which have been around for nearly 10 years, and URL shorteners.
http://example.com/?utm_source=June%2B2011%2BNewsletter&...
Source: June 2011 Newsletter
Medium: Email
Campaign Name: Free Summer Tickets
The moment that you store that URL in other service, a number of those tags become incorrect anyway (it is not an email anymore), and the stats you will get from it will be tainted. Campaign tags are useful, but this approach by pinboard may end up in tracking being more accurate (certainly from in terms of tracking campaign media/terms/content), at the cost of removing the campaign name.
There is also little difference between this and a URL such as http://example.com/Free-Summer-Tickets/June-2011-Newsletter?... being set up to serve the original content other than at least with the campaign tags you've got a single canonical URL using the more correct query parameter mechanism.
I also think, maybe I'm missing something. Maybe it's actually decent information to have. For instance, if you send a link with source=newsletter to somebody, it still is the newsletter that brought both of you to the site. You may not have visited otherwise. And your friend, probably even less so.
I don't know. I still don't like seeing it. It really does defeat the purpose of making your site have pretty links.
It sounds like you're probably misunderstanding the actual use case involved.
If you're a business, paying for traffic, you want to know ROI. If you spend X dollars and get Y visitors which make you Z dollars, you want Z > X, and you want to know what the relationship between the three values are, to know if spending more would be worthwhile. You don't actually care if the human being who clicked the link was currently in their email, rss, or anything else. You wan to know "spending dollars this way resulted in this profit". You want to identify the source, in your dollars, of your income.
http://news.ycombinator.com/item?id=2643515 The link contains this URL fragment: #.TfMwNJgETxs;hackernews
I think it's bad design to not assume that people can and will always share URLs with others. Putting tracking strings in URLs ignores this fact.
I often see Urchin URL parameters even on submissions on Hacker News. Most of the time they say utm_source=feedburner, which means that they were taken from an RSS reader. Just think about how easily such a submission reaching the top of Hacker News would distort the statistics.
Doesn't this make the the statistics almost worthless? Or at least harder to interpret? I think so.
There was a time where your site would be
site.com/product.php?id=x or
site.com/product.asp?id=x,
now it's:
site.com/keyword-keyword/keyword/keyword/product-name?urchin or google analytics crud
site.com/dashboard/alerts
We should get back to that. There need to be better ways to do SEO than polluting URLs. Slapping this information on a URL is a misuse of the purpose of URLs (permanent locators for a resource on the web. E.g. even if the resource has moved, there is a response code and a way of reaching the new location).
Example: http://www.w3.org/Provider/Style/URI
All the parameters you see specify campaign variables to help marketers track their campaigns. As for privacy, these parameters don't reveal user behavior on a site. It's only when it's connected to Google Analytics that is placed ON the site that the campaign data is then connected to site data. Long URLs have no adverse affect on the user browsing experience and URL shorteners do a great job of hiding them so removing those parameters simply serves to screw over marketing people that have spent time and money crafting their campaigns and gathering valuable data.
In regards to making it harder to detect duplicate URLs...really??? It's that hard to strip our everything before parameters?
it seems like a better way of doing it would be to capture those tokens at an initial url, but then redirect the user to the proper, clean url without them. that way you get accurate stats for visitors-from-rss-feeds, but everyone else that clicks through as that clean url is passed along appears as a different source.
url cruft is much more an aesthic thing. see http://www.mattcutts.com/blog/clean-up-extra-url-parameters-...
http://tools.ietf.org/html/rfc3986#section-6
I am currently working on making a urllib for python which is 3986 compliant
if (window.history && history.replaceState && location.search.match(/utm_/)) {
var check = setInterval(function () {
if (document.cookie.indexOf("__utmz=")!==-1) {
history.replaceState({}, "", location.pathname); //assuming you want no query string
clearInterval(check);
}
}, 500);
}"erode user privacy" - The data gathered is aggregated and doesn't identify individuals.
"make it more difficult to identify duplicate content" - You're already stripping them because you don't like them, so just ignore them when de-duping. Again, you can't have it both ways.
"and benefit ad publishers at the expense of everyone else." - You're out to screw the guys who pay for all those wonderful free services you use, like GMail for example:
;; ANSWER SECTION: pinboard.in. 3600 IN MX 1 s3.pinboard.in. pinboard.in. 3600 IN MX 2 s5.pinboard.in. pinboard.in. 3600 IN MX 4 ASPMX.L.GOOGLE.COM. pinboard.in. 3600 IN MX 5 ALT1.ASPMX.L.GOOGLE.COM. pinboard.in. 3600 IN MX 5 ALT2.ASPMX.L.GOOGLE.COM. pinboard.in. 3600 IN MX 10 ASPMX2.GOOGLEMAIL.COM. pinboard.in. 3600 IN MX 10 ASPMX3.GOOGLEMAIL.COM.
User-agent / IP / list of fonts that Flash can use pretty much identifies individuals uniquely. Adding "I got to this page by clicking a link on site X" to the URL adds one more piece of data that makes it even easier for the site to guess that you are you.
Can you think of a practical way in which this would actually make a difference?
My MX listings look awfully similar to that and it isn't a free service for me.