ClearURLs – automatically remove tracking elements from URLs
github.com
github.com
If you share a link from the TikTok app, it gives you a vm.tiktok.com/[xyz] link to send/post elsewhere. It gives you no indication that this isn't a generic link to the post, nor does it give you an option to expose the generic link to the post.
Instead, when you share that link and someone clicks on it and does not have the app, it opens with a header saying "[First Last] is on TikTok." On the other hand, once you do click on that link (if and only if you don't have the app installed), you get redirected to the static link to the video and finally obtain it.
This is an anti-pattern that enables further tracking and potentially unknowingly exposes user data when links are shared publicly. And there's no indication to the user that this is happening, since the link is structured as if it does not contain any tracking. Ie a tool like this wouldn't be able to "strip out" the tracking since it isn't tacked on in any way, but embedded as the generated link itself.
https://stratechery.com/2020/the-tiktok-war/
Any company running out of mainland China is going to have serious privacy problems due to CCP influence and their need to comply with both local laws and the government’s interest in influencing public sentiment.
I wrote about this at length here: https://zalberico.com/essay/2020/06/13/zoom-in-china.html and won’t rehash it again in the comments.
Yes, we know other countries have similar issues, but we can't excuse the blatant wrongdoings of the CCP by pointing the finger elsewhere.
It often feels like the work of bots or government shills anytime it happens, but good luck getting to the bottom of that.
>It often feels like the work of bots or government shills
Do you think I'm a bot because I disagree with you? Maybe you are the bot... how can we verify you're not? :D good luck getting to the bottom of that.
I believe China has kept tabs on 2 groups of non-Chinese citizens: 1. foreign nationals within China borders, and 2. foreign nationals who are ethnically Chinese.
If we're discussing the high cost of apples and someone brings up oranges, it doesn't change the fact that apples are expensive.
> Do you think I'm a bot because I disagree with you?
No, but a lazy comment doing s/China/USA/ certainly reads like it. And if you've seen some of the threads on Reddit or Twitter it becomes pretty clear some accounts search for any negative discussion about China and interject with whataboutism, which would be pretty easy to automate.
Yes it is, because the topic invariably includes "alternatives" or things that are "better". Those things are nearly always things the US has made. They are not better. Yes, it's bad. Everyone knows this. Adding value to that conversation is giving alternatives or offering some new insight into the nature of badness and how it's different flavours should be looked upon.
The issue with the China apologist comments is that their intent is to damage control by attacking the negative sentiments that would make the country look bad and steering the conversation towards the evils of other countries, i.e. whataboutism. This behavior is so prevalent online that you'd be hard pressed to find any discussion that criticizes China without it.
I'd also find it frustrating if in any thread that criticizes the US there would be comments about how China is worse. That might be the case, but it doesn't minimize US' problems and only distracts from discussing them. Yet such comments are much less frequent--I'm not sure if I've ever seen one.
Because it's painful to be awaken from the "American dream".
Assuming that this assertion is true what motivates China to be so authoritarian towards their citizens but not so to the rest of the world? Is it altruism or inability?
Does China only spy on their own citizens but not the rest of the world because they like the rest of the world more than their own citizens and they want the rest of the world to have rights and freedoms that they believe their own citizens don't deserve?
If it comes down to an inability to spy on the rest of the world what do you think will happen when China does develop the ability to spy on the rest of the world?
If you use Facebook or Instagram assume that the NSA has all your data, and that someone might try to manipulate you. If you use TikTok assume that China has all your data, and someone might try to manipulate you. You either choose your poison, or you stay on services that aren't in the limelight
I think the 'assume they have all of your data' is paranoid (particularly for encrypted stuff like whatsapp), but people should probably more careful about this kind of thing than they are anyway. The US has laws and rules around access, you may not agree with them - but they are far and away better than the CCP's approach.
The CCP is running concentration camps for a minority population of their own citizens, invading and taking over neighboring countries (HK with an eye towards Taiwan), and censoring pooh bear from the internet because of a light hearted comparison to Xi. The police call foreign students in the US to threaten them over their internet activity: https://www.vice.com/en/article/jgxdv7/chinese-police-are-vi...
The comparisons are not valid.
But as far as Google and Pakistan goes, most people who have an inkling of this think that the censorship only affects results served within Pakistan. But, in fact, the censorship affects search results served within the US. Google has allowed the Pakistani government, as well as various pressure groups and other governments, to influence what US people see within the US.
End-To-End encryption is useless if like in the case of WhatsApp you don't control the client, but a company beholden to US secret courts does. "“For the past decade, N.S.A. has led an aggressive, multipronged effort to break widely used Internet encryption technologies,” said a 2010 memo describing a briefing about N.S.A. accomplishments" [1]
> The US has laws and rules around access
I'm not a US citizen and reside outside the US, which from my limited legal understanding means that the US law doesn't give a crap about me
I agree that in recent decades China has a worse human rights record, which is a major factor when you "choose your poison".
1: https://www.propublica.org/article/the-nsas-secret-campaign-...
Even in Signal people can and do take screenshots, so really probably just best to be cautious of anything in writing that you wouldn't want published.
This is one reason I'm excited about Urbit - I think it'll be cool to get out of the dependence on centralized services.
Also, hopefully soon you will: https://urbit.org/
- IDs stop the spam problem and give people control over something that keeps its reputation (and they're cheap).
- Federated systems normally suck because administering the servers and keeping decentralized versions in sync is hard. Urbit's design fixes this.
- Encrypted by default, ability to be as easy to run as FB (eventually, not right now). Peer to peer with the address space and key issues solved from first principles.
- Stability over long time horizons due to design (goal being indefinite), the urbit abstraction layer doesn't change and state can always be recomputed - changing pieces are implemented via jets to communicate with whatever underlying OS is doing the normal stuff.
It's a clever design and solves a lot of problems with modern computing, people often dismiss it out of hand because Yarvin's politics are stupid (he's no longer involved in the project and hasn't been for some time). Peter Thiel's Trump support was stupid too, but that doesn't mean he doesn't get a lot of other stuff right.
You are correct that China cannot project power in the sense that they can't easily invade a country or level shattering economic sanctions but they have proven themselves quite capable of targetting individuals in other nations both online and in the real world.
Either way, there is a moral imperative to prevent China from gaining the ability to project power the way the US can. The US being able to project that kind of power is shitty, and the two entities being able to do that is even shittier.
I wonder how many non-Americans think two entities being able to do it is better than one, because at least they can counter-balance each other.
Not just with force; I was recently thinking about how the US during the cold war tried to be "nice" to the "third world" to keep them out of the "sphere of influence" of the Soviet Union. Currently China is trying to project it's "soft power" that way too to get less developed countries into it's patronage, but the US isn't really doing that at the moment (see for instance approaches to distributing covid vaccine...).
I (who is a usa citizen) personally am not really sure which is preferable, only one super-power, or two. Either way the world is in for a rough ride.
At least in the U.S., we have the intent of moral fiber permeating through our founding document, and probably at least half the population still fervently believes in these principles or tries to behave as though they matter.
China doesn't give a fuck. Their government is a communist dictatorship, and their sole concern is the expansion of their power through force.
Those who live in glass houses shouldn’t throw stones.
It's so hard to take you seriously when you make claims like this. Nobody cares about your founding documents, folks look at what the US does. And what it does is send drone strikes to hit schools, it bullies countries into doing what it wants without offering anything in return, etc. I genuinely cannot currently see a worse superpower.
With gratitude as a dominant force in ones's life, one is able to step back and see pictures other than what one wishes to see. One is able to stop buying into hysteria.
It is sad that some politicians use military as their play tool for nefarious purposes. Biden sent troops into Syria almost immediately upon assuming his current position. Reprehensible.
But... this doesn't apply to looking at China, only to looking at the US, for some reason?
This is a very confusing conversation, you seem to be switching the parameters of what we're talking about.
We started out talking about the general dangers of a superpower in the world, especially to people not citizens of that superpower. Is that still what we're talking about? Are the purported invention of air conditioning or the internet in the US relevant to that conversation? Is "gratitude as a dominant force in one's life" relevant to it? How does "gratitude as a dominant force in your life" effect your view of China? How should it effect the view of someone in a country getting significant investment or foreign aid (or cough vaccines) from China? Are you asking us to have a different attitude toward evaluating the danger of the US as a superpower to the rest of the planet vs evaluating the danger of China as a superpower to the rest of the planet? With one we should center gratitude and avoid hysteria, but with the other we should.... center hysteria and avoid gratitude?
Other countries don't use social media much, because they are culturally just not as interested in it.
There's a few countries where you can't really avoid being on any social network and that social network is not domestic, but those you can probably count on one hand. Off the top of my head I can just come up with Australia, India and Indonesia.
I do not agree with Chinese stance on democracy or human rights but I admire their willingness to play by their own rules. Not opening their markets and rolling out the red carpet for Silicon valley was wise.
I disagree. People will (somewhat) expect you to have a WhatsApp in many European countries, but hardly anyone will expect you to use Facebook.
At least in my circles Facebook is a wasteland. Many people haven't even posted anything in years, and if I was trying to reach anyone via Facebook I'd settle in for a long wait - until they check it in a month or two.
You won't notice it if you just open Facebook, because Facebook will fill your feeds with people who are active, but when I go through my list of contacts there it's obvious less than one in five are still actively using it.
I don't claim total knowledge of the situation everywhere, but I do keep in contact with people of a lot of different countries.
So mock and downvote all you want. I don't see why ClearURLs couldn't add this functionality.
Edit: Or am I just being downvoted by people who don't want anybody to know that it's possible to stop this form of tracking?
Or maybe you and OP don't care about the privacy part of the problem, and you just want to automate getting the "canonical" / "non-personal" one from the "masked" one?
You want ClearURLs (or something else) to always resolve to a canonical link, so that you're easily able to share this canonical URL, and to never have a tracked URL in your URL bar so that you don't share it by mistake. Makes sense.
But the average person isn't going to do that. They will share the nice, short, pretty url that tiktok gives them. But once someone gives you that shortened url, there is no way for you to view the video on the other end of that URL without being tracked. You would need to follow the link, tiktok would track you, only after they have logged the data will they send your browser a redirect to the proper url.
1. When I go to share a link, automatically trace it and remove all tracking so I get the final URL without any tracking parameters attached.
2. When I am sent a link with tracking parameters as a part of it, or a shortened link, send it to a remote server which will follow the links until it finds the final destination and removes tracking parameters, then send it back to me.
Both approaches have downsides. The first is nice for when I send a link to a friend but not when I get a link in an email from a company. This happens to me all the time and since I use NextDNS to block trackers I often can’t even get to the final website because of the various trackers I would have to go through to get to it which are blocked at the DNS level. I am still trying to figure out a good solution to this.
The second has the obvious privacy problem: who is watching the watchers?
ronjouch explains how it's not really possible to stop this form of tracking below. In order to unmask the URL, you need to pretty much visit the URL, which registers the tracking data, so even if you, as a user, gets a stripped URL that's safe to use, you will still have effectively clicked the link.
Unmasking (by the sender or a trusted intermediary, such as Tor) removes the risk of leaking the sender data to the (transitive) recipient
I think the problem is that, for security reasons, ClearURLs can't change URLs arbitratily. It can only remove parts of it, so the actual URL would have to be a parameter. See [1] for a relevant comment by the extension's author.
[1] https://github.com/ClearURLs/Addon/issues/102#issuecomment-8...
I think you are confused how this works. Because it would NOT be possible to stop this type of tracking. That is why you are being downvoted. The downvotes are because you are simply wrong, not because there is a conspiracy on HackerNews of people that don't want other people to know that it is possible to stop tracking.
Here's how it works: In the example given above, you only have the url vm.tiktok.com/[short-url-id]. This URL does not represent anything on its own. When you click the link, it goes to a tiktok server that looks up the `[short-url-id]` portion of the url in a database, which contains the actual video id/url that is trying to be shared, along with additional metadata about the share such as the person that shared it and the device the user is coming from, etc. This information is then logged in a data warehouse or sent down a data firehose to eventually perform advanced analytics to TikTok. All of this is happening while you are waiting to get the real url of the video back. Yes it's only a few milliseconds, but by the time you get the url of the video back so that you can actually watch the video, the data has already been logged. Your privacy is already compromised.
So your suggestion is to "unmask" the url and "untrackify" it and then give the user the end-url with the actual video. The problem is that the only way to get the real url and to "untrackify" it, you need to contact TikTok and they will already log the data before you can get the real url back. You can't simply "unmask" it. Only TikTok knows what the real URL is. In order to get the real url you need to ask them (by following the short url link) and they will log your data before they give you the real url. There isn't any way around this (other than not using the vm.tiktok share links).
I am not sure if the "real url" that tiktok gives you contains url parameters in it or not. It probably does. So you could theoretically remove those. For example turn tiktok.com/video-url?sharing_user=username123&device=iphone into tiktok.com/video-url. This would be possible. But it wouldn't do anything to protect your privacy. It would simply remove the "[First Last] is on TikTok" message. But the data already got logged when you exchanged the short-url for the long-url. So the privacy damage has already been done. This is why "unmasking" simply doesn't do anything other than give you the illusion of privacy, without any change to real privacy.
By contrast, when you see a url like cnn.com/news-story-url?utm_source=facebook and you remove the parameters from that type of link, you can actually overt a certain level of tracking because the tracking hasn't been logged yet when you remove the parameters. So removing the params into the link cnn.com/news-story-url and following that, will avoid the tracking because the tracking is done on the actual visit with that specific url. Since you removed the tracking parameters, the website now has no data to actually track.
Preemptively opening the link as the sender will send a request to TikTok, but they're not really gaining any useful data there since you just watched the video, hit share (this is what they know so far), and now you opened the link that you had generated. So their database only learned that you shared a video with yourself, which you immediately opened.
The more valuable data is when various intended recipients open the link, allowing TikTok to associate you with them to serve more targeted videos based on implicit social graph, etc.
Moreover, opening the link yourself to get the "canonical url" protects yourself if you're sharing the link broadly since others can't obtain your name [and potentially more?] from the shortlink.
Now, if you're the recipient, there's not much you can do to avoid the tracking link, besides opening it up in as much of an anonymous environment as possible. But interestingly enough, I find the privacy threat greater to the sender. The sender has a TikTok account to aggregate data quite straightforwardly, unlike the recipient. The sender is also being associated with a number of recipients, vs. the recipient with only one sender, and again only through cookies, IP, or something of that sort.
And that's actually stopping it. Even if you don't want to do that, there's real utility in an incremental step where if I go to re-share a Tiktok video I don't accidentally help them track others.
If any shortlink uses bitly as a backend, you can expand it yourself by copying the link and adding a "+" at the end, bringing you to the bitly properties page for that link.
For example, when I search Google for 'Hacker News', the URL I arrive at is "https://www.google.com/search?client=firefox-b-d&q=hacker+ne...". If I want to send that to a friend, I would edit the link to be "https://www.google.com/search?q=hacker+news".
The dedicated share buttons will often give you a link generated on the fly, with all the tracking info on the back end. For example, if google was to do this (which they thankfully don't), the link might look like "google.com/?query=cce1602b-5af6-4d95-965b-e88450afc266", and in the database there would be all sorts of tracking info tied to it. I can't edit that URL to dissaciate from that information, so if I share it, they would know it was me who shared it, and not someone else visiting it on their own.
Of course, companies can and do track you via less obvious means all the time, but this is just one small way you can foul a data point for them.
If they create custom urls for everyone that look like https://website.com/uuid/ and don't redirect you to the real url... it is not possible to strip anything unless you do some research to find another URL that redirects you to the same page. Not sure what that would do to your search engine rankings though...
A short video platform can hardly be expected to be a paragon of security and privacy. It has no utility whatsoever. I don't see where the concern comes from. A video of someone drinking coffee does not particularly invoke a point of concern.
What may be the real concern is China and the fact that the app is tied to it. Thats more race/geo-politics/war-mongering issue than a privacy concern.
Just like you wouldn't stand there listening to a drunk person complain about alcohol related health issues, I'm not about to entertain people complain about privacy when they have the agency and choice.
Ooo, that's pretty neat. I wonder if something similar can be achieved on Android. I usually manually paste it in chrome and copy the redirect, although I also enable desktop view to not get the mobile link.
* Accept a URL as input
* Expand the URL to the full link
* Find the "?" in the new url and snip everything after and including it.
Originally I looked for ".html?" but some TikTok links don't have the ".html" anymore so I had to switch to just "?". Tasker for Android [0] might be what you are looking for but I can't be sure. You might want to ask on the subreddit [1] for help or search there for something similar.
[0] https://play.google.com/store/apps/details?id=net.dinglisch....
I only noticed when I received a badge for how many times it was clicked, and even though it's not nefarious I'd still prefer it to be opt-in rather than done by default.
Bug 1697982: "Firefox Tracking Protection should protect against URL/queryparam-based tracking (like ClearURLs/NeatURL addons do)" , https://bugzilla.mozilla.org/show_bug.cgi?id=1697982
Please vote for the bug if you'd like it too.
Also, I see a few interesting comments in this HN thread; this evening when the dust settles, I'll aggregate & bring them to the bug for consideration if/when fixing this bug is considered.
and maybe a custom option; where you can toggle what is cleaned and whatnot.
and provide a "remember for this domain".
If utm_* query arguments are used solely for tracking, then it only makes sense that Firefox goes the next step
Good luck with that. They have no choice but to believe whatever data the browser sends them, data that we control. If their precious content leaves their server at all they've already lost.
This is a bit like antivirus software authors worrying about being "banned" by the virus creators.
Affiliate links are often hidden, and depending on the system might even lead to higher prices, because the shop is offsetting the affiliate program cost.
Whether the company does A/B testing, what does that have to do with me? That also can be implemented without external trackers and just be set in a session.
So I would say it’s a net-positive.
All attempts by Mozilla to bake-in addon-like behavior so we don't have to install 'yet another damn addon' is welcoming, but as with any of these features, they come with caveats already present in the addons.
For example, Firefox's HTTPS-Only mode (that is basically the HTTPS-Everywhere addon) breaks some sites, and also their anti-tracking feature will break some sites too. But then again: if a site is serving HTTP only then they're doing it wrong (with the exception of captive portals). As for the anti-tracking feature: I rarely see sites asking me to disable my AD-Blocker, and when I do I never give-in, no matter how desperate I am to see hidden content.
Every time I have used an add-on like ClearURLs, I have had issues at some point due to some zealous clean-up of URLs which breaks a redirection.
Typically, I don't want the browser to mess with my browsing if I am on the websites of my bank, a shop, etc.
1. First, by Mozilla analysts & developers making a good job at rolling out a potential implementation in a safe progressive way, with the easiest stuff first (`fbclid`, `gclid`, etc), and then going deeper / per-site, maybe re-using (part of) existing filterlists.
1.1. Also, note that ClearURLs is quite aggressive (as noted by a few commenters, and I confirm): it strips lots of non-URLbar requests, strips ETags, etc. A sibling comment mentions that alternative NeatURL is less aggressive. As with all cat-and-mouse games, this is a trade-off, and an implementation in core Firefox doesn't have to go as far as ClearURLs, at least initially. Offering a strictness knob to users is also an option.
2. Then, Firefox already has UI to disable Tracking Protection and work around sites broken by it: click the shield at the left of your URL bar, then toggle off "Enhanced Tracking Protection is ON for this site" to see if it was ETP that broke the site. This UI maybe need adjustments / more granularity (and maybe not), sure.
outgoing.prod.mozaws.net/v1/25c02fd4e609951729e0ec0b41fe5391d912511b45d2a02aeaa839872c8d9def/https%3A//gitlab.com/KevinRoebert/ClearUrlsAny URLs in the addon description section are all tracked/redirected via `https://outgoing.prod.mozaws.net`
I use the HTTPS only mode in Firefox - it breaks some sites, and telling Firefox to disable the mode for a specific site doesn't always work.
I feel like a plugin (HTTPS Everywhere) can deal with this a lot better than something that's integrated and reduced to a single checkbox in the settings.
(At this point, you or a passerby will point at the Pocket fiasco and argue that there's too much stuff shoved into our browsers and just stahp it already. Fair, and I love lean software too. I'd still like this specific feature because A. it's not Pocket, B. it aligns well with what Firefox is doing these days, and C. it aligns with what I expect from my user agent of choice).
Then, supposing this ever makes its way into Fx, you can choose not to use it. And by the way, maybe like you, I will make the same choice if the Fx feature is too basic! But it will remain a win, for the users for whom it's good enough and who would never have bothered with an addon in the first place. Just like ETP vs. uBlock / PrivacyBadger / etc: ETP is a good basic "80%/20%" risk-less step in the right direction, and addons remains way ahead if you the user decide to bother a bit more.
> "telling Firefox to disable [HTTPS only] mode for a specific site doesn't always work."
This looks like a bug that you should report.
And Stop trying to "re-implement" features for which there are already user extensions way more capable
(and then the tracking economy, let’s say China for example, will just steamroll our economies. This is what I‘m worried about, in a vacuum a slower developing but ad/tracking-free society would be preferable of course.)
Of course I despise all ads as much as the next hacker here on HN, I just wonder sometimes if they‘re a necessary evil.
So in the end I‘m inclined to agree with your nuanced „some general statistical gathering is OK, just no fingerprinting etc“.
Such an economy is still very possible: just pay people for their data. Giving it away isn't economically efficient and imposes significant negative externalities, as we've seen.
Oh well. Just let the economy perform slightly worse then.
> for knowledge means better resource allocation
Who cares about some corporation's resource allocation? That's their problem to solve. We should be caring about all the people whose privacy they are violating instead.
If they want to allocate resources efficiently, they should be required to do it in a way that doesn't invade anyone's privacy. If that means they'll make less money so be it.
(I've noticed that TikTok does this explicitly, providing a different short URL for each share request - it's clean, which makes them easier to share without blocking out a whole chat, but still not wanted)
It's not that I don't mind the parameters, it's that I also mind the URL tracking. And I can do something directly about the URL tracking.
They are fundamentally first-party analytics - they show up in the server logs of the site visited, and that site needed to craft the link in order to place the parameters in the first place. There's a big difference between URL parameters and e.g. cookies attached to third-party javascript.
I definitely support the freedom of people to remove these URL parameters if they want. But it's not fair to classify them as a "war" - they are a tool used by scrupulous marketers, too.
Even if there are scrupulous entities, the harm caused by the unscrupulous ones overshadows them.
In the case of some large sites that provide pervasive services like FB, Twitter, and Google, you interact with their sites incidentally as you surf in the internet. It's these sites that are a potential privacy risk IMO.
And on and on it goes.
If you think about it... what's the problem with a URL tracking which advertiser you came from? Why would you insist that it be a secret which ad you clicked to come to a site?
This tool is removing URL parameters some of which are absolutely harmless and not violating anyone's "privacy". We really need to draw the line somewhere and decide what the heck means "privacy" at this point, because everything can be interpreted as violation of privacy.
Likewise, are those site owners allowed to exist, or should they just offer content at a loss, and pay millions of ads, and have no even clue which ads worked and which didn't? And when there's a paywall of course everyone is SUPER ANNOYED by the paywall.
So to recap, the public wants absolutely everything, for free, and they want to disrupt as much as possible from the site's mechanism to understand what the other side of this communication is and what they want.
Good.
> what's the problem with a URL tracking which advertiser you came from?
It's additional bits of information used to identify me.
> We really need to draw the line somewhere and decide what the heck means "privacy" at this point, because everything can be interpreted as violation of privacy.
Okay. If I explicitly give you information and you use it for my benefit alone, it's not a violation of privacy. Everything else is.
Concrete example: people provide their addresses to companies so they can have packages delivered. This is obviously legitimate. Selling my address to marketers so they can spam my inbox with unwanted ads is obviously unacceptable.
Placing identifying information in URLs is unacceptable simply because I didn't explicitly choose to reveal that information. I don't even care if it's harmless, the sheer audacity of these people is offensive.
> This tool is removing URL parameters some of which are absolutely harmless and not violating anyone's "privacy".
Yeah, I'm not risking it. They'll probably find a way to abuse this information if they haven't already. Marketers are not supposed to get any data whatsoever. I'm increasingly convinced marketing shouldn't even exist to begin with.
> should they just offer content at a loss, and pay millions of ads, and have no even clue which ads worked and which didn't?
Don't pay for ads in the first place.
> And when there's a paywall of course everyone is SUPER ANNOYED by the paywall.
That's okay.
> the public wants absolutely everything, for free, and they want to disrupt as much as possible from the site's mechanism to understand what the other side of this communication is and what they want.
I guess. Just return 402 Payment Required if people are expected to pay. We refuse to be the product.
Is the information you collect unusable for directly or indirectly identifying an individual? Go wild! :)
Facebook can use it to link contacts together. I get a share link, it gives it an ID, I send it to someone, they open it and now they have linked my account with their account. Same works if I click on a page and get the ID, share just that page, and someone clicks it (and there's some fb element on the page).
Now if several users a day share a link here on HN, facebook will know about us as belonging to a certain group.
This is different than the type of tracking you're talking about. ID-type parameters like gclid, dclid, and fbclid are all unique to the ad impression, and tied to the individual that ad was served to. Which means they can tie it back to other data sources they have about the individual. Like social graph data for Facebook, or demographic or interests data for other advertiser networks.
Personally I care about ID parameters a lot, and UTM-type parameters not at all. But that's just me.
https://aliexpress.com/item/4000336900709.html?spm=a2g01.126...
vs
https://www.aliexpress.com/item/4000336900709.html
Both take you to the same page.
URLs, especially when clicked from ads tend to have a HUGE amount of extra crap that's in no way needed for any kind of functionality.
The first link is:
https://aliexpress.com/item/4000336900709.html?spm=a2g01.12617084.fdpcl001.8.2665mpnYmpnYMH&gps-id=5547572&scm=1007.19201.130907.0&scm_id=1007.19201.130907.0&scm-url=1007.19201.130907.0&pvid=65430901-7ec6-4584-a620-4618974e03ae
The second: https://www.aliexpress.com/item/4000336900709.html
(HN normally truncates links to ... 60 characters, it seems.It's not fun to get a 2048 character long URL to a product linked to me on mobile, where it literally doesn't fit on the screen all at once. =)
I don't. I believe marketers should have exactly zero ways to measure the effectiveness of their mind hacking efforts. Any data they try and collect should have negative value by virtue of being completely randomized by the browser.
Actually I believe marketers shouldn't even exist. Nothing they say is trustworthy by virtue of conflict of interest. The internet would be much better off without these constant attempts to subvert it for their purposes.
Nothing they say is trustworthy by virtue of conflict of interest
Everyone who says anything has that same conflict of interest. You do, I do, marketers do, salespeople do, engineers do, politicians do, scientists do. Completely dismissing value of an entire profession based upon self interest doesn't have a limiting principle.
Marketing, even if you naively limit the term to just cover advertising, is a rich and useful function of capitalism and society in general. The key to dealing with it is in protecting basic freedoms like a right to privacy.
I do because in 99% of cases it's a deliberate waste of my time and attention.
> Everyone who says anything has that same conflict of interest.
I don't think so. The information I receive from friends and peers is far more trustworthy. With marketing, I get selective truths at best.
Lots and lots of people on this site admit to adding "reddit" to their searches when looking for product reviews. Why? Because they don't trust marketers. We want real information from real people with real experiences, not some paid-for narrative. We especially want to know the risks, the negatives and the cons, precisely the kind of information marketers want to bury.
Since ads need to be conspecious by law, I don't see a conflict of interest. We know this is a carefully crafted story of the person who has stakes in the product or service.
Think of your favorite site with the best experience possible. That is possible because people tested countless times what works, what didn't, what is the most efficient path to a rewarding UX, and so on.
Yes, there are a ton of garbage lazy marketers in the world. Saying that marketing shouldn't exist would immediately render every refined UX you have navigated, purchased from, and or loyally stream content from.
Throwing out the good because of the bad is too far of a reach IMO. Anywho, that's just little old me and my opinion doesn't mean much.
Just accruing swaths of data doesn’t help, you need to interpret it correctly. I think qualitative data will bring you a long way. Once you need to do A/B testing, you can also do it privacy friendly.
If you market your product and run a campaign? Why not offer discount codes or something to figure out how you got them.
Tracking what button or page layout works better from a conversion perspective is not a privacy issue. It's a user experience benefit.
Having a SaaS business and not understanding the exact user funnel, conversion, abandonment, etc. will directly translate into a loss of your job and/or the failure of your business.
This isn't about personal preference which you have every right to. This is about building a business, which is why we're all here, and understanding how to successfully delight our customers.
Funny; those kinds of sites are my least favorite. All those colors and buttons are an information overload, and the animations make my laptop fans spin like crazy. Not everyone bought their computer under a decade ago.
Please, blue links and black text aren't evil. We need to make interfaces functional and stop rather than continuously A/B test them to maximize addictiveness ("engagement").
"Think of your favorite site with the best experience possible."
Regardless of the site experience that you prefer, I can assure you that thought, testing, and iterations have occurred to deliver the experience that you personally prefer.
Furthermore, I don't want hyper-optimized experiences. These experienced tend to be addictive, whether intended or not. Using an interface shouldn't feel "magical", it should work. I know that you weren't implying addictiveness or engagement, yet these values are (consciously or unconsciously) prevalent enough in the field that they've lowered my level of trust in analytics-driven iteration. Other responses in this thread should also show that I'm far from the only one who feels this way. Earning back broken trust in these situations typically requires going above and beyond past expectations.
User research is research, and should require informed consent held to the same standards of consent as actual research. In any human research, participation should be opt-in. Participants should be given complete information about analytics and how they will be used (with the option to see source code), own their data, be able to revoke their data, and see conclusions of the studies in a format they can understand. Any questions they have should be responded to before and after they opt-in. This shouldn't be buried in a confusing privacy policy but provided upfront, in a language they can speak. People frequently learn interfaces in languages that are foreign to them, but reading details of user research is a different story: you might need a translator. Otherwise, your sample will be even more heavily biased.
This is a lot of work, and might make analytics more trouble than they're worth.
Good riddance. Affiliate schemes just encourage people to spam low quality content full of affiliate links to products that are rarely good.
And once the adtech companies notice that every tag except utm is being stripped, you can bet that utms will start being stuffed with tracking.
startswith: 'utm_', 'ga_', 'hmb_', 'ic_', 'fb_', 'pd_rd', 'ref_', 'share_', 'client_', 'service_'
or has: '$/ref@amazon.', '.tsrc', 'ICID', '_xtd', '_encoding@amazon.', '_hsenc', '_openstat', 'ab', 'action_object_map', 'action_ref_map', 'action_type_map', 'amp', 'arc404', 'affil', 'affiliate', 'app_id', 'awc', 'bfsplash', 'bftwuk', 'campaign', 'camp', 'cip', 'cmp', 'CMP', 'cmpid', 'curator', 'cvid@bing.com', 'efg', 'ei@google.', 'fbclid', 'fbplay', 'feature@youtube.com', 'feedName', 'feedType', 'form@bing.com', 'forYou', 'fsrc', 'ftcamp', 'ga_campaign', 'ga_content', 'ga_medium', 'ga_place', 'ga_source', 'ga_term', 'gi', 'gclid@youtube.com', 'gs_l', 'gws_rd@google.', 'igshid', 'instanceId', 'instanceid', 'kw@youtube.com', 'maca', 'mbid', 'mkt_tok', 'mod', 'ncid', 'ocid', 'offer', 'origin', 'partner','pq@bing.com', 'print', 'printable', 'psc@amazon.', 'qs@bing.com', 'rebelltitem', 'ref', 'referer', 'referrer', 'rss', 'ru', 'sc@bing.com', 'scrolla', 'sei@google.', 'sh', 'share', 'sk@bing.com', 'source', 'sp@bing.com', 'sref', 'srnd', 'supported_service_name', 'tag', 'taid', 'time_continue', 'tsrc', 'twsrc', 'twcamp', 'twclid', 'tweetembed', 'twterm', 'twgr', 'utm', 'ved@google.', 'via', 'xid', 'yclid', 'yptr'
Edit: Will turn this into a Gist at some point.
In your second list, are those the names of query params? I'm puzzled by the inclusion of @ in many of them, maybe you're saying that '_encoding' is a tracking param on any amazon domain, 'sk' is a tracking param on bing.com? What does the $ in the first entry indicate?
https://github.com/jparise/chrome-utm-stripper/blob/0d16a13d...
Before
https://example.com/item/4000336900709?spm=a2g01.126...
After:
https://example.com/item/4000336900709000044323234
Now the tracking parameters are all encoded in the last segment of the url. The backend just has to decode it accordingly and it will have both the item id and the bag of tracking parameters.
So something would look like:
https://externalwebsite.example/some-article
but would link to
1. Find "rel=canonical" in the page's source.
2. Look up the page/article title on your favorite search engine.
> The extension can read the content of any web page you visit as well as data you enter into those web pages, such as usernames and passwords.
I'm sure the devs are super trustworthy, but there have been cases of legitimate extensions falling in the wrong hands, and this, coupled with automatic extension updates, could be a big security hole in your setup.
[0]: https://support.mozilla.org/en-US/kb/permission-request-mess...
PS: Ironically, the link above has utm elements.
Discussion about ClearURLs permissions: https://gitlab.com/KevinRoebert/ClearUrls/-/issues/159
Extension permissions: https://developer.mozilla.org/en-US/docs/Mozilla/Add-ons/Web...
Besides that I only have the Tree Style Tab add-on installed, which is much recommended.
Isn't uBlock Origin a better extension, both from a blocking (no "Acceptable Ads", for example) and also performance point of view?
Since you have Firefox, you could sync with a community-developed user.js like Arkenfox (previously GHacks) [1], which seems to go much farther and still not break much! At least the settings privacy.resistFingerprinting and privacy.firstparty.isolate looked indispensable as soon as I learned what they do.
And without FPI (first party isolation), not getting LocalCDN [2] (Decentraleyes successor) and Temporary Containers [3] seems like a gross oversight. They have a great discussion on add-ons at the Arkenfox wiki [4].
[1] https://github.com/arkenfox/user.js
[2] https://addons.mozilla.org/en-US/firefox/addon/localcdn-fork...
[3] https://addons.mozilla.org/en-US/firefox/addon/temporary-con...
Fingerprinting is extremely hard to avoid, without being painfully conformist.
But fair point about "extreme" vs "sane". They're quite subjective terms.
I use Chromium, which is mostly plugin free, as the alternative.
While the ETag header may have been usable for cross site tracking at some point in the past [1], browser caches are isolated per-origin in Firefox, so there's no longer a cross-site tracking concern. That leaves it usable to identify you across sessions only in a first-party context, just like cookies, IP addresses (to a lesser extent), the Last-Modified header, and any number of other identification techniques ClearURLs doesn't block.
[1] I'd be interested to see any credible evidence of ETag headers being used for tracking in the wild - I've only seen theorizing that it _could_ be used as such, prior to cache isolation being implemented in Firefox and Chrome.
> ETags can be used to track unique users, as HTTP cookies are increasingly being deleted by privacy-aware users. In July 2011, Ashkan Soltani and a team of researchers at UC Berkeley reported that a number of websites, including Hulu, were using ETags for tracking purposes. Hulu and KISSmetrics have both ceased "respawning" as of 29 July 2011, as KISSmetrics and over 20 of its clients are facing a class-action lawsuit over the use of "undeletable" tracking cookies partially involving the use of ETags.
It appears that there have been at least a few cases of this in the wild.
The main distinction (at least to me) between ETag and the other tracking methods you mention is that ETag doesn't appear to be easily clearable by a user (although that sounds like something browsers should fix if they haven't already).
It's unfortunate that features like this end up getting co-opted by trackers, which leads to breaking legitimate use cases like your app in the process.
The Last-Modified header can be used in exactly the same way, and isn't blocked by this extension, which harkens back to my original point: this is an extension that appears to see significant use by non-technical users, yet it breaks a browser feature by default. There are plenty of other methods of identifying a unique user that it doesn't prevent, so this seems like a pretty unexpected feature users should take note of.
Sites put this in because they want search engines to index a single clean URL rather than many tracking URLs, so it’s pretty reliable.
javascript:window.location=window.location.href.replace(/\?([^#]*)/,function(_,s){s=s.split('&').filter(function(v){return(!/^utm_/.test(v))}).join('&');return(s?'?'+s:'')});
It's much limited as it focuses on Google's links, but it works good enough for many cases.One things I noticed is that it can be too aggressive from time to time. I encountered this "issue" when creating a Bitwarden account, I was unable to verify my e-mail address because ClearURLs was (unbeknownst to me) removing some of the parameters from the activation URL. While similar cases will most likely not be frequent, it can be really frustrating to determine why something does not work (also applies to ad blockers).
I'm tired of every time I want to share a product page or post a URL or something, of having to strip 300 friggin' nonsense characters from the end of it.
Open Source and Free to Use
Cookies is one way (if we stick to the domain), possibly using sidecar AJAX requests and localState is another.
Or maybe we can leave it all in the URL, but encrypt it with a key in a cookie, thus without the cookie, the info is recognized as foreign when passed around. Hmm yeah, not bad.
A caveat: When you submit a request to archive a url, archive.is sends the client-ip (X-Forwarded-For) to the destination server.
archive.is archives webpages by stripping it off of any dynamic content, but it tries its best to capture a dynamic (multi) page (spa) anyway (for example, twitter threads, linkedin profiles). It limits file-sizes up to 50MB (I think?), and works best with text-heavy webpages (news and blogs, for example). archive.is is ran by a for-profit company based in NY, but it isn't clear who is in fact behind it. Ref: https://en.wikipedia.org/wiki/wp:archive.is
I already see many sites use something like ?arg={BASE64 STRING OF ALL THE THINGS} and no automatic tool can decypher that as it's a custom list of bytes.
> no automatic tool can decypher that
...
But yeah, home grown analytics can't be reliable circumvented.
Didn't know uBlock Origin has this, giving it a try. Thanks ChrisGranger!
For now I know of two lists from prominent maintainers which purpose is to remove unneeded URL parameters:
- https://filters.adtidy.org/android/filters/17.txt
- https://raw.githubusercontent.com/DandelionSprout/adfilt/mas...
There are ongoing discussions to include the first one as a stock list (i.e. present in "Filter lists"), though not enabled by default for now.
Addendum: to be clear, this is not a replacement for ClearURLs. ClearURLs has more capabilities then just removing query parameters from the URLs of outgoing network requests.
---
[1] https://github.com/uBlockOrigin/uBlock-issues/issues/1356#is...
One by-the-way question: how can I "discover" such 3rd-party lists? Is there a place doing lists centralization / aggregation / recommendation?
docs.python.org$removeparam=highlightThis effort is really ought to be shared, it's potentially a lot of manual work, and could benefit many projects. ClearURLs seems like one of the most promising existing projects doing similar stuff; have been meaning to approach the devs, feels like it's something we could cooperate on. Although ClearURL has a somewhat narrower scope, but still I feel like there is a potential to share.
I'm also thinking that it might be possible by some simple machine learning, by looking at the corpus of existing URLs. E.g. if a human looks at a corpus of different URLs they would more or less guess what is useful, and what's tracking garbage, so perhaps it's possible to automate it with a high accuracy?
Then, I also feel if it's paired with some UI to allow the user to 'fix' the algorithm for entity extraction (e.g. by pointing at the 'relevant' parts of the URL), it would already be good enough for the user -- they would fix the sites that are worst offenders for them. Then these fixes could be optionally contributed back and merged to the upstream 'rules database'.
I assume this extension doesn't deal with that problem (redirect-type tracking URLs).
With Apple's Universal Clipboard, you can clean links copied on your iPhone
Publishers are desperate to monetize their audience anyway possible. Affiliate revenue always seemed to be lesser of evils, IMO, in comparison to programmatic/display. After all, the user intentionally is making a purchase vs. having their data sold out from under them with zero knowledge.
Here's to hoping that I'm misunderstanding how inclusive this will be to stripping parameters.
https://techcrunch.com/2021/05/04/fewcents-raises-1-6m-to-he...
If you use any site that's broken by the extension, that means you have to remember to turn it off (globally!) before using the site, and turn it on again after.
That said, it still results in some level of tracking and given the add-on's purpose, having to opt in to affiliate links seems like the right choice.
It seems the interesting bit (for me) is in the Rules[2] repo.
[1] https://gist.github.com/yb66/d39109df620ab1db2a46c943111c31d...
https://apps.apple.com/us/app/resolver-share-clean-urls/id14...
In addition to cleaning trackers it tries to convert AMP and Apple News links to the real pages.
Or you can manually add the filter now. After installing the Ublock Origin addon in firefox mobile, I clicked on the 3 dots -> Addons -> Ublock Origin -> Open the dashboard -> Filter lists -> Import... and pasted this URL (from the top of the above link): https://raw.githubusercontent.com/AdguardTeam/FiltersRegistr...
I tested by sharing a URL with UTM and other parameters, and it did strip them.
I'm glad that uBlock is, everytime I browse with Chrome, which doesn't have extensiona at all, it seems like a dystopia.
> Prevents Google from rewriting the search results (to include tracking elements)
How infuraiting. Now sites open without my having to pay google 1s-2s for them to log me.
Great extension idea.
If this one is better, any chance supporting it in ublock?
Privacy is a right. Ad-tracking is not.