The mounting cost of stale ad blocking rules (2018)
brave.com
brave.com
1. Use Selenium and the DevTools Protocol to record every URL requested when rendering and executing a website.
2. Add additional automation to randomly select three distinct same-domain URLs from anchor tags on a page.
3. Used the above automation to visit the homepage of each site, and a maximum of three child pages, and recorded all URLs requested for images, script files, and other web resources.
4. Determine which of those URLs would be blocked by the version of EasyList fetched on that day, using Brave's optimized ad-block implementation.
...
We found that the vast majority of EasyList rules are not used when browsing popular websites; 3,268 of 39,198 (~8%) of network and exception rules were used during our crawls (these measurements exclude element rules)."
That doesn't mean that EasyList is not useful for browsing the rest of the internet.
The idea here would be to compile the regex of each rule to an automata (each automata matches a set of strings) and them build a fst out of them (that would map strings to matched rules)
Matching against static, glob and regex patterns. And there is some grouping and filtering so, it's not like every URL gets all 64k patterns matched. And it's possible to pre-compile and group the Regex to get a little faster (lots of framework routers do this)
From my opinion the "web" is so contaminated it's a shorter accept-list than a constantly changing reject-list
Furthermore the decision tree needn't be binary. If you group regexes by common substrings and then even group substrings for, say, a Commentz-Walter matcher, you are able to distinguish N+1 cases for N patterns, with significantly lesser costs than N times the cost of matching for one substring alone.
Arguably its more useful to know what a given website actualy requires than to know every possible advertising scheme. Allow lists teach about where to find desired content. Block lists do not teach anything useful in that regard.
A hash table can't be used for this kind of check because it uses patterns, so resources need to be compared on a pattern. Although I'm sure there are special cases, like hash tables of domain names, which cover a large portion rules.