‘Just four dudes’: Inside EasyList, a community-run ad-blocking list
digiday.com
digiday.com
My takeaway was that it's pretty impressive they still respond to exclusion requests in a timely manner.
Maybe it's good just to have a reminder that much of our infrastructure is maintained by volunteers. Sometimes important policy decisions get made by whoever showed up and started doing the work in a way that looks competent to others.
Articles are supposed to have a point, a goal, etc. The goal might be persuasive, it might be to establish facts, it might be to argue for a specific change or against one. It's up to the reader to agree or disagree with what the article has laid out.
But we need something to agree or disagree with. In this case, I had no idea what the point of the article was (beyond a vague feeling that the author doesn't like EasyList).
RewriteRule "^/+new_path/+(.+)$" /advert/$1 [L]
... for each "new url path" (as, granted, they had quite a few which matched "advert".... but that's also something they could again fix with another RewriteRule, I guess?
RewriteRule "^(.*)/+new_path/+(.+)$" $1/advert/$2 [L]
... would've worked just as well?If parts of your audience blocks ads then it's probably a good idea to test your web site with uBlock installed, just as you already test with Safari and MSIE.
It's specifically not a good idea to rely on the block list maintainers who will have a hard time separating earnest bug reports from all the bogus ones who think they've found a clever way to avoid blocklists.
Even so, you could fix everything with a one line Perl command.
Please.
list maintainers should be responsible regression test their lists for normal content and not do dumb things
this is also a problem with bad/unscrupulous spam listings which list domains based on unsubstantiated requests, etc
I am against advertising and marketing and yet spend a considerable amount of money online, and I’m confident most of HN does the same.
They responded and removed amp-analytics. I'm not that impressed.
I worked on an early version of a big woodworking website and it was over six thousand static html pages.
The group that acquired it needed to show progress asap, which meant simply updating the header to their new logo and footer nav.
I had to write about a dozen sequenced find and replace routines with regex to do even basic changes. In some cases groups of pages used subtly different html it was a real mess.
It took a long time for them to get all the pages replicated in a templating system to avoid seo loss.
https://forums.lanik.us/viewtopic.php?f=64&t=43148&sid=e0edf...
My task was to move all of that into a very minimal CMS. All the content would be in flat files, and the page itself would load the header, footer, content, etc.
Because every page was different, but mostly similar, I wrote an Emacs extension that let you identify the content areas (it tried some heuristics and you could guide it) and it would do the extraction and emit the flat content file and the replacement JSP file.
What you have to realize is that many people have systems that are designed for 1 "thing" and have since "scaled" that up 1000x. People still treat their website as a directory of arbitrary HTML. People manage 1000 servers like they managed their one and only very first server. It is the natural way of growth. Engineering is figuring out when that model doesn't apply anymore, and changing it. Engineering is in short supply, however, leading to stories like mine above or the "thousands of pages" blocked by EasyList.
I wouldn't be at all surprised if AutoTrader used manual editing for everything, with no access to scriptability.
Not everyone is using jekyll or any modern stacks ... the echo chamber is strong here. I guarantee you 65% of the Internet is still being edited by hand.
I once worked on a site that had over 25k pages, all hand coded HTML with Microsoft Frontpage. And all in the root folder.
Not even the owner knew how many pages there actually were.
> "It’s crazy that more people don’t know about this," said Marty Kratky-Katz, the founder of Blockthrough, which lets publishers monetize using Adblock Plus’s Acceptable Ads program.
If only there were a business service that might help me navigate this perilousness of which I have just learned, so that I might better monetize.
It makes me giddy to consider how much the quality of the internet will improve when the bottom-feeders are at last excluded from it.
The same issue was seen with Spamhaus and similar IP block services for eMail spam. What were selfless intentions quickly morphed into ineffective, power-crazed gatekeepers and extortion artists that have made SMTP an unusable protocol unless you are Google or Amazon.
I started working on http://blockedby.com , a monitoring tool to warn you when a rule would impact your website.
But after contacting a few websites that did have the issue on the forum, I couldn't find any that actually wanted to prevent this. A few discussions with publishers told me that they're more interested in convincing their users to give up on ad blockers rather than face the fact that people are not going to abandon ad blockers.
I'll be taking a closer look at the tool.
I'm no fan of obtrusive, JavaScript laden ads, and have used adblockers myself for years, but why block "legitimate" content?
I didn't know how to get removed from the list (and assumed it wouldn't happen anyway), so changed a couple of class names and it was working again. But I don't know how long the issue was there before I realised it :/
Anyway. People rarely intend to block legitimate content. At the same time, it's not always easy to tell what is and isn't legitimate when you're writing fairly general rules.
`<a href="our_twitter_page"><img src="/img/locally_hosted_image.png"></a>`
Was everything hosted on your domain there? If they have images hosted by twitter/facebook that's enough for them to track users which is what these lists don't like.
`<a href="our_twitter_page"><img src="/img/locally_hosted_image.png"></a>`
The block rules were purely based on very generic CSS class names - I forget exactly, but something like "social-twitter", "social-facebook"
It's now part of a bigger "annoyances" list so maybe that was included somewhere in one of the filterlists.
Seems a bit heavy-handed to me personally. I absolutely want to block intrusive ads and tracking, but this was a bit OTT. It also must have been included by default with uBlock Origin, as I didn't add it myself.
I realized the blunt force approach of various adblockers will block anything on "/ad/...". Some would even block "/ad...". I found that a bit annoying.
So, was it simple <a> elements linking to your Twitter/Facebook page?
Or was it maybe a metric ton of obfuscated JS fetched from a Twitter/Facebook CDN, that would relay every bit of data about your visitor, and as a side-effect display a button?
My take after reading this, is that the author wants EasyList to either not exist or to be run by "experts in balancing publisher monetization" (aka adman)
Just another day of unsavory corporations and the assholes that run them trying to mislead the average consumer.
Sometimes reading about things like "balancing publisher monetization" puts the picture in my head of a bunch of anthropomorphized tapeworms trying to have a public policy debate.
They were elected by people voluntarily installing EasyList. We had enough of you experts™, thanks.
Blockers like Ublock Origin also don't benefit from legitimate sites breaking; if they're turning on EasyList by default it's because they think it's currently the best balance available for their users.
If anyone can make a better list than Easylist, then they should just go do it. In fact, companies already tried to make an alternative with the Acceptable Ads initiative -- and if their list is better, people will switch to it. The only barrier of entry to displacing EasyList here is quality over time.
TIL According to Wikipedia: "In July 2018, uBlock.org was acquired by AdBlock"[1]
This is not uBlock Origin.
https://github.com/gorhill/uBlock/wiki/uBlock-Origin-is-comp...
It's true that it's about a year out of date, but I've never had any ads get through regardless.
On both Google and DDG ublock.org is still the first match when you search for "ublock". I wonder how many people install this crap extension by mistake every day.
Imagine if a company launched a (non-Google) "Chromium" browser or non-Microsoft "Explorer" or "Edge", how would the authentic developer defend against that? (Probably with a trademark suit, which is a dificult strategy for a tiny open source project).
gorhill could create a brand new name to replace Origin, but hat runs the risk of losing even more people to the knockoff "uBlock".
Perhaps "uBlock Official" would be a better name than "uBlock Origin", but ultimately it comes down to a sog of hard marketing work to protect and build mindshare.
Maybe fork uBlock Origin, give it a new name, update the current uBlock Origin official pages to point to the new one, keeping links in older threads still relevant for a while.
I'd do my part and promote the new name as much as possible. I'm already spending a lot of time typing out the explanation about the current name whenever I get a chance.
I’m pretty sure Eyeo did register the uBlock trademark, and has a good case to sue for someone implying through their product’s name that it’s “the official” uBlock.
Yeah they are, the balance they're targeting is zero ads.
Good, very good. Break them all. Learn to go back to basics, use TABLE, not gazzilion of nested DIV's for what should be a single BUTTON element instead. I run uBlock Origins, NoScript and Privacy Badger; and whenever I go, to clients, friends, family, I'll always put them up. Your shitty site has the same merchandise like many others and if your "buy" button is broke I'll just go to the next site...and while on this, how come Amazon "buy" button is not broke? Almost feels like Amazon devs actually test with ad-blockers their content before allowing it on the wild.
Yes! I'm doing this too, as a kind of community service. If sites get broken because their privacy-invading "like" button doesn't work, that's on them. People will just move on.
Fandom (aka Wikia), in particular, has run some incredibly intrusive -- and occasionally even malicious -- ads on some of their sites. I have very little sympathy for them.
No, for big sites the adblock users report issues, or more likely the adblock list developers just fix the issues right away. The issue is small sites with no sizeable adblock-using population.
FYI those four dudes are whats keeping me from disabling javascript and images from every website on the face of the earth.
Thank you EasyList maintainers!
Why would you ever want your adblocking list run by "experts in balancing publisher monetization"?
However, I suspect that Easylist already supports acceptable ads in the form of "route the ads through the publisher's domain/servers" which puts accountability where it belongs, so it doesn't need further expert advice.
Maybe your moral fiber. Definitely not mine.
If I'm getting something that required skill and labor, 'moral fiber' suggests some quid pro quo. How are you going to reciprocate?
At some point someone observed "we can profit off of audiences" so that answer was "by being an audience member." But now it's kind of a mess, and either we haven't found a better model or inertia and collusion have prevented it from arriving.
On a tangentially related note, I think there's a special place in hell for the engineers that created gas pumps that display ads while fueling. Those things are pure, unadulterated evil.
But we’re talking about moral fiber. For some people “they hit me first” is not ethically sound. There’s some ambiguity here about who hit first, but rather than a fair exchange we are getting what is rapidly approaching brinksmanship. Nobody is really in the right here.
But we also can’t seem to just go back to a paid model for high production values. We are satisfied with amateur work and underpaid professionals slowly burning a nest egg and hoping something changes. The advertisers are the only ones with a hand out, so they get to set the agenda.
I had hopes Brave was going to alter this trajectory but there seems to be some regulatory capture there. i kind of wonder if they hired some ad people and charisma won over the origin story.
Why?
I've got a lot of questions I'd like to ask about why exactly that's a big task.
And also they should talk about uBlock Origin. Not AdBlock Plus or UBlock.
It seems even plausible, reading it, that the author knows about uBlock origin but avoids naming it, instead listing the two sellouts that are morally shady.
Fine people like Raymond Hill will keep fighting the good fight but I wonder if someday we'll look back on the era where web content was paid for by the 90 percent of people who don't block ads so the 10 percent could block them effortlessly as a sort of golden age.
Unrequested audio is flat-out absurd behavior in a browser.
In a sense, it's already kinda happening with Twitch. A few months ago, they rolled out Server Side Ad Insertion [1] which broke ad blockers. Now, ublock origin and streamlink can still block them though but the arms race will continue. I don't think it'll ever come to reencoding the video to incorporate ads because most ads are dynamics but maybe future encoders or hardware are really fast.
[1] https://aws.amazon.com/media/tech/what-server-side-ad-insert...
Yes! I think what has staved it off is the lack of adblocking on mobile devices. I use uBlock on every computer I have, but on my iPhone? Not so much. I think as more traffic moved to mobile it has offset ad blocking to some degree. I can't imagine that 25% of people are blocking ads on mobile?
If enough sites started using anti-adblock techniques with toxic ads, my next step would be a 'site blocker' plugin that would effectively eradicate it from the internet for me. This is kind of the nuclear option. No links from others sites, no showing up in search results. The technology exists already for blocking adult sites: the thing missing is the ability to add to your block list from a search result page or on the site itself.
The next logical extension to that is sharing block lists among users, and projects that maintain block lists you can subscribe to.
It actually would be nice if there's a way to notify the site when you were blocking them, such as your browser hitting `/.well-known/eradicated?origin=http://page/containing/link`.
Maybe we'll get a resurgence of text-only websites.
I would argue that too much of the content that is paid for is created purely to generate page-views to carry advertising.
Those people who are motivated to share knowledge and opinions sincerely are crowded out by click-bait.
We'll need some clever strategies to counter the arms race.
Image side-by-side visual diffing. Some process renders pages with and without ads. Ads are progressively identified and removed until it renders like the "print" or "reader" view.
Temporal diffing. Snap shot popular websites over time. The meat (content) will likely remain the same. Everything else is chrome or ads.
My other ad detecting notions are even more harebrained, so I'll stop there.
We discovered a few years ago that our list of advertisers (in a digital magazine) was being blocked by ad blockers. It was coming over the wire with the word "advertisers" in the URL. These aren't ads that pop up and display on the pages, but simply a list of advertisers in the issue.
Rather than try to get adblockers not to block it, we simply used another word in our URL and went on with things. We shouldn't have had to. Those are lazy adblock rules and are incredibly likely to get false positives and break sites.
But we're realists and knew that even if we fixed the existing adblockers, another would come along and do the same lazy thing.
This claim comes accross as disingenuous. These guys are trying to block ads accross the whole internet without ruining the user's core experience. Maintaining a hard-coded black/whitelist of URLs and elements simply wont cut it. Sometimes there will be false positives.
If this impacts enough of your users, then it's your responsibility as a web developer to test your site on environments that reflect your users' and make something that works on their environment. You can always choose not to support those users, but that's your problem.
"Just four dudes" sounds dismissive, but everyone starts out as "just one dude" (or dudette, or whatever the proper female form of dude is), and sometimes, more join. Often that's not the case, and the project's health hinges on just one maintainer.
Even if these pages are static files (although any site opetating at that scale should have some kind of page-generation system), this still sounds like it could be resolved with a simple find-replace using <insert your favorite text editor>.
EasyList is definitely in the right here. Adding an exception results in an unnecessary performance penalty and creates more complexity for the project. Both of these impacts are small, but they can't afford to accommodate for every website like that.
If they want their users to use their site, they should make it work on their user's environments. I doubt they would demand Google or Mozilla accommodate for them like that.
Reading the article, "balancing publisher monetization" provoked an immediate visceral vocal response from me. Truly hilarious.
They couldn't be widely shared, but they're less whack-a-mole in practice and from a security perspective, a more superior solution.