AdGuard publishes a list of 6K+ trackers abusing the CNAME cloaking technique
github.com
github.com
The CNAME of the Game: Large-scale Analysis of DNS-based Tracking Evasion - https://news.ycombinator.com/item?id=26347110 - March 2021 (51 comments)
If you haven't tried out AdGuard Home, I can highly recommend it. Has same feature set as Pi-Hole and support DoT as well. It's also super trivial to install since it's just a Go binary. Have been using it for ages now and love it!
I actually discovered it when I wanted to install pi-hole on my mac server and it just wouldn't work besides with the Docker container, which had other issues like not being able to see the client IP that made the request.
Been running AdGuard Home for a couple months now and it's really nice!
I have encountered multiple times (not common, but not trivial) that a filter blocks something wrong. With traditional ad-blocker as extension, I can quickly find it out by using build-in logger, and then simply either temporarily disable them or add the site to whitelist with a single click (if I feel like it, I can write my own rule too.)
If I have to change my DNS setting everytime this happens with these DNS-based blockers, I feel like to stick with extensions since I don't really use my phone to browse Internet too much.
The only problem is browsers like Chrome that are pretty aggressive with DNS caching.
I found that, after tinkering with blocklists for a bit, I turned off logging altogether and just let it run. The one thing that gives us grief occasionally is (unsurprisingly) tracking links from promo emails and social media. These are usually easy enough to bypass, but it can be a pain for non-tech-savvy people.
Does it have "cosmetic filters" (the ones that block certain elements on page) or similar feature?
AGH already supports adding AdGuard filters, but for obvious reasons it only applies domain based filters. Adding the MITM proxy would allow for processing of the cosmetic filters too.
More info:
edit: I almost wonder if this could be done safely in a decentralized manner. Everyone runs a BitTorrent like service and when you make a request the swarm proxies it to a specific node who then serves the page back to you.
Under this scheme it would become impossible to trace who you are.
I guess the only thing that’s necessary now is to set it up so that the node could be another mobile client. Maybe it’s possible with webrtc or an extension?
Even if it's encrypted, I feel like someone smarter than me would figure out how to do bad things with this.
Personally I don't think the strength of p2p is only trust, it's delegation. If you have peer to peer encryption running, the possibilities are endless.
Add a statistical "proof of authenticity" and you have an unbeatable anti censorship mechanism that can also identify in-page modifications and weed out malicious MITMs in between.
in the future, given sufficient CPU resources on the clients, another method could be browser page-rendering engines that use advanced machine learning image recognition to categorize and blank out ads, no matter where they come from.
[1] https://github.com/bslassey/ip-blindness/blob/master/near_pa...
I think Google pre-emptively locked down Chrome API to prevent ML based blocking from working. This would explain their fixed URL block list and removing interactive API
https://proxy.example.com/ZGQgaWY9L2Rldi91cmFuZG9tIGJzPTEgY291bnQ9NTIgfCBiYXNlNjQgCg==
However, the "end game" of adblocking is far worse - the entire page becomes this: <!doctype html>
<html lang=en>
<head>
<meta charset=utf-8>
<title>null</title>
</head>
<body>
<canvas>Run the WebAssembly blob to render the page</canvas>
<script src="load_webasm_blob.js"></script>
</body>
</html>
The entire page becomes an giant obfuscated WebAssembly blob that renders the page into the canvas tag. The technologies of the open web like HTML become legacy baggage; a "web page" is just a stub loader for what would effectively be a statically linked executable binary that uses the canvas tag as a generic framebuffer.In this "end game", URL based blocking is irrelevant; most page assets become part of the blob. The question "is this an ad/tracker/virus?" becomes undecidable; answering any question a page's behavior requires running potentially hostile code or solving the Halting Problem. Some people don't want an open web that respects things like user agency. They want control. They want the "web page" to be an opaque binary blob that nobody can investigate or modify. They want full control over what the user sees and is allowed to do. They want TV.
Of course, the publishers will need to trust the advertisers, but I don't see how they have much choice...
> Under this scheme it would become impossible to trace who you are.
This is how i2p works, in a nutshell. Every i2p user is a node in the network for others to proxy through.
Unfortunately, hiding your IP and geographic location isn't enough to stop fingerprinting and other forms of de-anonymization.
/img/logo.png?uuid=..&res=1920&os=MacOS&osv=11.2
and when heuristics catch up with that ...
/img/logo.png?757569643D2E2E267265733D31393230266F733D4D61634F53266F
ie, se09d.png
Also used sometimes for standard cache invalidation efforts so common for both nontracking purposes too.
Worst case sites will just install a middleware layer to proxy all their requests and traffic back to ad tech machine I'd imagine or install some middleware on their stack.
Think along these lines: website.com/blog/f5/babBys-first-bl0g-post-2f.html
Can you (or a computer program) answer definitively whether there's a tracking ID in that URL?
But these changes are still good for privacy. These direct-proxy endgame methods will hopefully make it harder for ad companies to detect ad-fraud, making the ads less economical to begin with.
It's not like publishers will give up on integrating third-party tracking code into their tech stacks once this technique is dead. They'll have to do so at a deeper level, e.g. reverse proxies, NPM modules, etc.
Next battle in this war: adtech code becomes indistinguishable; tracking IDs start appearing in all exit links from major publishers; adtech firms & publishers coordinate on the backend to sync those IDs. What then, get rid of URLs?
that can be easily filtered out eg. https://addons.mozilla.org/en-US/firefox/addon/link-cleaner/
- Loss of HIPAA and PCI compliance.
- Loss of trade secret protection for contents of the website.
- Liability for security breaches due to third party capture of sensitive information.
And anything else that follows. If e-commerce sites with dubious trackers and ad networks can’t take credit card payments, they’ll quit the trackers.
Even if track.a.com and track.b.com are both CNAMEd to eviltracker.com, the server at eviltracker.com doesn't have any special ability to crossreference traffic between those two domains. At this point, you're effectively blocking first party tracking.
Except, if you make it hard for sites to employ even SaaS-based first-party tracking solutions, then you probably encourage them to roll their own or use hosted first-party tracking, which is less likely to support features like cookie opt-outs or GDPR compliant data handling.
Sure it does. It generates the IDs in the first place, so it knows how if they match between the sites. Not perfectly precisely (at least not for now), but with high confidence by the source of traffic and browser fingerprint.
> use hosted first-party tracking, which is less likely to support features like cookie opt-outs or GDPR compliant data handling
On the other hand it doesn't aggregate data between unrelated services. And the first party can access a lot of the same data just from standard logs.
It's just a hack around analytics software needing an on-prem deployment or server-to-server integration. Whatever!
But it's still exfiltrated data, and if the user was given the choice they would almost certainly choose to block such requests.
Further, it strikes me that such trackers engaging in this behavior are knowingly subverting user wishes. Which is sort of demonstrating bad faith. I believe attempting to demonstrate the bad faith of tracking companies was a motivator for the DNT flag. They slithered out of that issue by arguing that because people were setting it as default, users didn't effectively consent to not being tracked. This case seems more clear, where they are subverting anti tracking plugins.
There's no cross-site aggregation (much less in-browser tracking) possible in this model, unless I'm missing something major.
As long as you can run scripts, you can be matched between sites. But even before you get into completely unique identifiers, you can get lots of data from repeating ip/location - people are fairly predictable that way.