EasyList is in trouble and so are many ad blockers
adguard.com
adguard.com
But what about simply enable a firewall and show captcha or similar if the origin IP is from India and requesting that URL until the situation is under control? I did that with the free plan recently in CloudFlare in a similar situation and it worked perfectly (of course on a much smaller scale).
If so, returning a bogus file that blocks everything and adding a comment in that list asking the developers to use caching or mirroring the file should be fine.
I wonder if those browsers honor the list when fetching the update though. Would be awesome if you could just add easylist and lock out further requests right on the device.
They just need to serve either an empty response or an intentionally broken rule to break the misbehaving browser and force its developers to fix it.
However what you can do is match the user-agents, and return a global/catch-all adblocking rule that blocks all the content of all the pages (by blocking the body element).
The app developers are going to notice the issue very fast (because users are reporting the problem), and mirroring the lists or adding a cache is immediately going to be their priority.
Bonus: I think some browsers and extensions can execute JavaScript in adblocking rules; https://help.eyeo.com/adblockplus/snippet-filters-tutorial
(which is essentially re-using a gigantic XSS in order to notify the user)
I read that poisoning your own lunch to catch a workplace fridge thief could be considered assault.
EDIT: here’s what I read. https://law.stackexchange.com/questions/966/can-one-be-liabl...
Imagine, say, you update the list to block all URLs, and it impacts some municipal government worker’s ability to update some emergency alert service and causes hundreds of people to be permanently injured.
Same for Easylist, if they decide that a quota of 100000 requests per IP+UA per day is the maximum, that's their choice. They owe nothing to the consumers of the lists.
That being said; Easylist actually benefits from being distributed in many apps; it is really valuable to influence / control adblocking lists, so the more flexible they are to the browser developers, the better (I guess).
Very much in the same way that image host use to change an image for those hotlinking directly to images in the early days of the net.
Providing a service (which you expect others to consume) and then not only deciding to refrain from providing, but "poisoning" the output, is an interesting move. We don't consider them equivalent, but in a case where this application was providing some essential service that is not easily replaced, and physical harm was a result, how do we consider it?
The made up scenerio of this preventing some critical task from being accomplished is stretching at best.
However, as we’re living in the real world and the authors of the respective browsers strike me as lazy or uninterested, I also bet all that would change is the user agent.
Because they certainly want to serve some huge amount of traffic for free while they attempt to become the next abusive monopoly platform.
They’re trying to have their cake and eat it too.
Cloudflare provides a significant service to the free and open web by subsidizing the hosting costs of static content for websites. They give that away for free under what appears to be reasonable terms.
But the comments here about Cloudflare’s ToS read a lot like folks feeling entitled to getting bandwidth for for free. Cloudflare is providing a very specific service for free, and it does a lot of good.
It would be great if Cloudflare decided to donate. But I’d re-evaluate your stance if you’re feeling entitled to their resources.
Rather what's being condemned is this nonsense customer service characterization of a text file as somehow not "web content". Easylist.txt is a data file that could just as easily be in JSON (and be larger). Furthermore, as it stands easylist.txt actually looks like it's a valid text/html file, as browsers generally don't insist on <html>/<body> tags. So from both directions it seems like the customer service drone has thrown out this nonsense just to short circuit having to do their job.
And there is no reason a web app couldn't read data from a txt file instead of a json file.
This is a very straightforward interpretation of the terms and it's strange to see such pushback based on a pedantic technicality when it's clear what the file is being used for.
In fact, if this was the opposite situation and some automated rule was involved in isolating this file, I expect the same people would then want human intervention to clear up the difference based on the context.
And that's just as nonsensical. Bytes are bytes; the rationale should be based on bandwidth, not on arbitrary micromanagement of the format of the data consuming that bandwidth. If I encode that video in a giant self-contained blob of JavaScript that feeds the pixels into a canvas or something similarly ridiculous, does that magically fix the bandwidth issues?
Again this is that "pedantic technicality" - why such a fuss when the actual issue is straightforward, and also clearly understood and reiterated by the easylist team themselves in the post?
Then the policy would focus on that rather than micromanage the format of the data using those resources/bandwidth. Again: bytes are bytes.
> why such a fuss
Because a policy as nonsensical as "no non-HTML files allowed" artificially limits the usefulness of CloudFlare for precisely zero legitimate reason. I ask again: does wrapping a video in a blob of JavaScript fix the bandwidth issues associated with hosting videos? If I have a 10MB MP3 downloaded 1,000 times v. a 1MB HTML/CSS/JS static site downloaded 10,000 times, what difference does it make?
Also, CF have a product to sell. The free tier is just the demo version: I think at the end of the day the policy is about not everyone in HN using CF for their low-cost DIY video and/or music streaming or download platform.
And I can totally see them reverse the decision and sponsor that project (it's probably something a support engineer has no power to decide).
Also, the same ToS applies if you use the paid product. Unless you buy an additional addon for non-web traffic.
Okay, now run the same thought experiment with 10,000 downloads of a 1MB MP3 v. 10,000 downloads of a 1MB HTML/CSS/JS site. What difference then?
> why such a fuss
Mockery isn't a fuss. And why? Because we can all picture having to deal with just this sort of idiot and their pet rule.
The majority of of legitimate traffic to these txt files are by browsers(extensions) for the express purpose of displaying websites to a users specifications(without ads in this case).
A reasonably low estimate is that 20%-30% of global internet users are behind some kind of adblocker, almost all of which are default subscribed to easylists. So this txt file is potentially responsible for the way a BILLION+ internet users see and interact with near every single website.
Cloudflares claim that this isn't web content is full on reality warping. It only makes any kind of sense when masked under layers of abstractions and lawyer speak.
---
None this should even matter though, someone seeing the big picture at CF should have done the napkin math and realized that the easylist bandwidth pays for itself.
Ignoring the soft cushion of the huge amount of globally distributed caching of these files, if easylist suddenly stopped working for a week then global bandwidth usage could see a spike. A pretty little chunk of which CF may be on the hook to absorb at no cost.
It literally is - specifically, for extensions thereof.
Clearly EasyList lived on their free tier for a long time without interruption. Only when they used excessive bandwidth did ToS enforcement happen. When they reached out for support, the support agent rightly pointed out that this isn't a website file.
Reading the ToS, the support agents message appears to be correct. Text files are fine (as is pretty much any format) as long as it isn't the main focus of the HTTP server Cloudflare is fronting. Robots.txt would be fine, turning the list into XML or HTML would not be fine. In this case, the text file isn't there to support the web content of Easy List - it's distributing a text file to applications.
The agent could have added additional context but their message is valid.
Meanwhile -
1. easylist.txt is used by every single web page I visit. So the overall purpose argument fails.
2. Web pages commonly use non-directly-renderable data files in formats like JSON or XML, so the file purpose argument fails.
3. Text files are and continue to be one of the major formats displayed by browsers. So the file type argument fails.
4. The size of the file is in line with other files cached by Cloudflare. So that argument fails.
If the Cloudflare support rep said "we just don't feel like doing business with you", that would be a different thing. But instead they're throwing out some arbitrarily-framed unfalsifiable reason as if it's a logical justification. And no, customer service drones and corporate policies don't deserve a fundamental benefit of the doubt, per contra proferentem (ambiguous terms should be construed against the drafter). It's impossible to know what they actually mean here besides "we don't like it", and that is the problem.
But in order to make the support eng do his job I need to add a .html extension to it?? Would that be considered a website file?
You're missing the point. If Cloudflare's issue is with bandwidth, then they should say so and leave it at that, not conjure up this pathetic excuse about .txt files somehow not being "web content". Does wrapping that data in <html><body><pre> </pre></body></html> magically fix the bandwidth issues?
All you'd need is a <pre> tags or <style> somewhere if you'd want it rendered not just as one large paragraph.
I guess you may need a <!DOCTYPE html>
I'm not sure this is the most reasonable rule but there are definitely some benificial aspects to it. For example the load on human-viewed content is limited by how often people want to view it. Not how often their browser wants to redownload it.
By Cloudflare's rationale it does.
> For example the load on human-viewed content is limited by how often people want to view it. Not how often their browser wants to redownload it.
Bandwidth is bandwidth. If 100,000,000 humans want to download a 10KB text/html page v. 100,000,000 programs wanting to download a 10KB text/plain file, both within the same time period, then that's going to be the same degree of load on Cloudflare's end.
Many websites also push text through HTML as part of AJAXy stuff. If they actually enforced this for all sites, their service would no longer be usable.
https://developers.google.com/search/docs/crawling-indexing/...
Still, I'm sure the 'community' can figure out how to keep something like this online. I'd be happy to pony up some cash for decent hosting and I'm sure many would be. If that doesn't work out, something like ipfs, a torrent or whatever.
> A great opportunity right now for CloudFlare to win some goodwill and PR by helping out EasyList for free right now.
> EasyList tried to reach out to CloudFlare support, but the latter said they could not help. Moreover, serving EasyList actually may violate the CloudFlare ToS.
Seeing the comments from Cloudflare here, looks like the HN machine has yet again worked its magic to get appropriate attention!
Unfortunately, I can't find a reference to it anymore.
https://en.wikipedia.org/wiki/Censorship_of_Wikipedia:
“Wikipedia has been blocked in China since 23 April 2019”
⇒ putting ads for Wikipedia on sites likely isn’t safe everywhere.
I think it will be very hard to find “some other organization” that is universally ‘approved’ everywhere.
The thing is, they didn't host the scrapped images themselves, they just hot-linked everything.
So through a little nginx config, we turned their entire homepage to an ad for my friend's platform :)
In earlier days of the Web, someone appeared to have hotlinked a photo from a page of mine, as their avatar/signature in some Web forum for another country, and it was eating up way too much bandwidth for my little site.
I handled this in an annoyed and ill-informed way, but which I thought was good-natured, and years later realized it was potentially harmful. I'd changed the URL to serve a new version of the image, to which I'd overlaid text with progressive political slogans relevant to their country. (Thinking I was making a statement to the person about the political issues, and that it would be just a small joke for them, before they changed their avatar/signature to stop hotlinking my bandwidth.) Years later, once I had a bit more understanding of the world, I realized that was very ignorant and cavalier of me, and might've caused serious government or social trouble for the person.
Sensitized by my earlier mistake, I could imagine ways that a subtly NSFW image could cause problems, especially in the workplace, and in some other cultures/countries.
Subtle things like flipping the image upside down or reversing the colors or other "not quite harmful but quite annoying" responses are probably better, or just serve a 1x1 pixel image of nothing.
https://news-ltn-com-tw.translate.goog/news/world/breakingne... (nsfw)
https://www.rfa.org/english/news/china/japan-piracy-09252022...
1: https://news-infoseek-co-jp.translate.goog/article/president... (og: https://news.infoseek.co.jp/article/president_61325/ )
2: https://i.imgur.com/5hjqu3L.jpg (label on bottle and window sign)
https://www.vice.com/en/article/bvxb94/is-this-beverly-hills...
I was getting hotlinked from controversial sites a lot at one stage, and the common forum software they used didn't force image sizes. So a 5k pixel wide image pushed most of the content off the screen thanks to a centred element :)
This was done for 2 reasons:
1- Avoid scenarios like this where you ship code (extension in this case) that is hard to update. Then make that code depend on external resources outside of your control.
2- Leak our users' IP addresses to each random hosting provider.
So the solution was simple: Run a CRON once a day then host the files ourselves. Pretty happy with that decision now.
It’s great that they’re happy with their choices, but the choices would, in this same situation, likely saddle them with a crippled infrastructure and/or some insane bandwidth bills for suddenly pushing 100 extra TB/m.
Edit: so if OP wasn’t implying that their approach was better, what was the point of posting it? Wow, obtuse much?
Edit:
OP was implying their approach was better for themselves than relying on third party servers. It’s hardly obtuse, it’s a related discussion from someone who would otherwise be impacted by the throttling.
OP, in contrast, wrote their own ad blocker targeting their own servers. They're in control of their ad blocker code and can write it to be respectful of their servers. They're not hosting the lists with the intent of allowing other people to use it, and they're unlikely to attract lazy app developers because the endpoints are (presumably) not listed publicly on the internet to anyone who wants an easy ad blocker list.
On download.easylist.com have it shuffle and send a redirect to a mirror to download the list. I wonder if Universities are still offering these small amounts of space for Open Source projects?
So robots.txt is not supported by Cloudflare to cache/proxy it? That would be a weird regulation. And I bet everyone violates the Cloudflare ToS then.
it's cloudflare deciding to protect "web content" and not videos or .iso images or other things that normally are not commonly served while you browse a contemporary website and read HTML.
That's false in two ways: first, text is normally served while you browse a contemporary website; second, so are images, which are explicitly called out as potentially violating this clause. Text is the only data that isn't covered by this clause.
It's all 1s and 0s too
2.8 Limitation on Serving Non-HTML Content
...Use of the Services for serving video or a disproportionate percentage of pictures, audio files, or other non-HTML content is prohibited, unless purchased separately...
A huge text/plain artifact, requested often, would seem to fall into that category of "disproportionate percentage" compared to text/html served.
But anyway, just rename .txt to .html and you're done.
At least file extension is limited and externally visible (and thus accountable) to third party behavior, which should limit the worst complexity excesses.
Is filesystem metadata actually different (theoretically) from extension? Or just data in a different format?
Extension seems a nice balance between simplicity / brevity and utility, albeit as a hint, not a commandment.
Cloudflare just seems to be trying to limit the free tier to "caching website html for the purpose of showing it to humans". They have pricing and plans for things other than that.
What else is important to note that the client is being abused and not the client abusing the service. That should be taken into consideration, when deciding if someone is breaking the ToS.
> What else is important to note that the client is being abused and not the client abusing the service. That should be taken into consideration, when deciding if someone is breaking the ToS.
My understanding has this as moot. The issue from Cloudflare's perspective is only that the content is non-HTML and doesn't have anything to do with the rate of traffic (the abuse).
The key is "as viewed through a web browser" imo, this is not really an API and it's not a webpage; it's a datafile and would fall into R2 or similar things.
Cloudflare is balancing on a razer for this TOS technicality.
To oversimplify, they’re saying Cloudflare’s service is to be used for serving websites to browsers.
Serving a static text file that is primarily used by applications is not in line with their terms of service.
Cloudflare provides a significant service to the free and open web by subsidizing the hosting costs of static content for websites. They give that away for free under what appears to be reasonable terms. I’m not sure why you’re trying to “gotcha” through their ToS.
It would be great if Cloudflare would donate resources to EasyList - it would do a lot to help the free and open internet by giving users more power over what gets delivered to their browser. But call that what it is: a donation.
People are doing the opposite, pointing out the hole and asking them to get a better rule. Surely they don't just want the list merely converted into html.
> They give that away for free [...]
So they should specify things that influence cost such as total bytes served, number of files, etc. Currently all you can do it bypass the rule because you don't know how to cooperate.
Imagine you do that and I DDoS the URL. CF will then mitigate this DDoS by, in part, replacing your html with their Browser Integrity Check html.
If you're serving 'web pages and websites' everything continues to work. What would happen if this list suddenly became an actual webpage.
If your site is serving 'a disproportionate percentage' of non-html you decrease the ability of CF to tell good traffic from bad.
<!DOCTYPE html>
<title>a</title>
Practically, browsers will accept omitting both of these, and the spec even allows for omitting the title "if it is provided by a higher level protocol"So it's not that crazy an argument that a plain text file is a html document
They serve websites to browsers for people to view. This file (be it properly formatted .html or .txt) is not a website people go to in their browser - its used internally by an application. This is the key point.
May be EasyList could host them there? That's what we do [1] (and the dashboards show 400TB+ per mo [2], likely rigged by the traffic between Workers and Cloudflare Cache).
[0] https://news.ycombinator.com/item?id=20791660
[1] https://news.ycombinator.com/item?id=30034547
[2] https://nitter.net/rethinkdns/status/1546232186554417152
text/plain though is decidedly not text/html and I would expect CloudFlare to potentially do some on-the- fly optimizations that are aware of the structure of an html file that save terabytes a day at their scale.
Some think its very Oracle of Cloudflare to do so. I do not blame them.
I mean, one could simply wrap the content in a HTML body and change the extension, but that would actually increase the data load for no good reason. So it is complete non-sense to complain about txt files being served.
> Use of the Services for serving video or a disproportionate percentage of pictures, audio files, or other non-HTML content is prohibited, unless purchased separately as part of a Paid Service or expressly allowed under our Supplemental Terms for a specific Service.
We will never know the reasoning of the support agent who replied to the EasyList maintainers, but I can imagine that it is indeed disproportionate for EasyList.
I really hope that Cloudflare actually sees that they are making a wrong decision here and actually help the EasyList maintainers.
Though, adblocking is a big business, many actors there are getting large revenue.
For example, Eyeo's income was 50 million USD per year last time I checked (and I guess most of it is actually profit), so they can find a solution if they really want.
https://img.phantasmagoria.me/img/96XJrjejoHNdrQv7.jpg
Even if you have a private bucket, you can give people a signed link with read access, for up to two weeks, IIRC.
But it'll still cost them money by number of reads
I assume you probably know this but just wanting to share there are some pricing scales with R2 they're just pretty generous for a lot of things.
If they're downloading text you can still use the headers, and some tricks around redirects, but overall you have far less data on which to decide.
Even more wtf- the file extension determines the file content?
> Based on the URL that are being requested at Cloudflare, it violates our ToS as well. All the requests are txt file extension which isn't a web content
> you cannot use Cloudflare to cache or proxy the request to these text files
This does not mean banned.
They didn’t ban the whole organization, but effectively told them to stop. I don’t see a difference.
CF didn't do this. They sent them an email telling them that what they were doing was a violation of their TOS and to cease doing it. They did not kill off their account. They still have the option to comply and continue with CF, which seems to be what they are going to do at the moment.
Hopefully, CF will grant them amnesty on this one. At the end of the day, an HTML file is just a text file, so I don't see why this would have even mattered to begin with.
Do you have a source for that? The article only mentions them being throttled + the screenshot with the support engineer saying they seem to be breaking the ToS and asking them politely to move back into compliance.
Google built an entire browser and used Manifest V3 as an excuse to cripple ad blockers.
Companies are also paying influencers, twitch streamers, and YouTubers to promote their products in a way that conventional ad blockers can't prevent.
The very interesting thing is that none of Google's ads have ever made it through this new version of Ublock for me.
Which I'm okay with in the same sense that I'm okay with newspaper/magazine ads or billboards or TV/radio commercials: they're annoying, but easy to ignore compared to online ads chewing up CPU time and battery life while actively violating one's privacy.
Yet. One day someone will create an ad blocker with machine learning that "sees" the ads and deletes them in real time. Should work on all content types, even on augmented reality.
Also, when doing auto-updates: always add a chaotic delay offset 1 to 180 minutes to distribute the traffic loads. Even in an office with 16 hosts or more this is recommended practice to prevent cheap routers hitting limits. Another interesting trend, is magnet/torrent being used for cryptographic-signed commercial package file distribution.
Free API keys are sometimes a necessary evil... as sometimes service abuse is not accidental.
At this point, they might be better off coordinating with the other major adblocker providers and just outright move the file elsewhere. Breaking other people's garbage code is better than breaking yourself trying to fix it. Especially on a budget of $0.00.
If the defective code for the browsers are in public repos, it might also be more effective for someone to just fork the code, fix the issue (i.e. only download this file once a month, instead of every startup), and at least give the maintainers a chance to merge the fix back in.
to
https://127.0.0.1/file.csv?apikey=abc123
This could allow client specific quotas, and easy adoption with maintained projects in minutes. Thus, defective and out-of-maintenance projects would need manually updated or get a 404.
=)
In this case, it would need to be distributed to myriad users who legitimately need to ask for the lists and then could be scraped by the "attacker", but at least then they'd have to be knowingly malicious vs. accidentally malicious.
Then browser makes like this will not reasonable be able to request a new key automatically for every install. So they will just request one and ship it.
Then when you get abuse like this you can disable it.
In general, it is easier to setup filters after differentiating legitimate from nuisance traffic. For example, fail2ban looks at the log of errors from invalid hits, and bans the IP or entire ISP block ranges for 5 days. This ban than get propagated to the rest of the web via spamhaus.org listing.
i.e. the users start to see the worlds internet become unreachable, as admins start to block traffic at the NOC's routers, and so on... India knows about Karma.
I am more surprised the app store for the apk isn't getting sued for theft of service.
I will also point out that if you fail2ban the IPs requesting the spamblock list, it could become worse if the browser just retries endlessly in the background. The traffic for a 404 page could be much smaller than the traffic of the very same devices trying again every few seconds, constantly, instead of only checking that 404 every app restart.
As a side note, some people build spider traps that reply with a pre-baked bzip file as a spoofed HTML compressed response. Thus a client program dutifully decompresses a few TB sized document, and browser exits due to memory issues. Note most modern Browsers are wise to this trick these days, but I doubt a dodgy plugin disk-usage limit check would catch a client side storage-flood. People shouldn't do this though, even if it is epically funny and harmless. =)
On your side note; EasyList probably does not want people to start suing for starting to distribute malicious content to users (and crashing your browser on purpose is arguably malicious).
1. take the top 200 most popular websites in the given nuisance area
2. add ban rules to a version-B list that also includes all social media, search engines, and Wikipedia.
3. Look at the user-agent string for that specific problem client, or extreme apikey quota abuse
4. Randomly serve version-B filter list that breaks the browsing experience after a frequent update. Increase random breakage until traffic rolls off to normal levels.
The TOS for the ban list file does not specify which sites it will ban, and most users will just assume it is the App that is broken (it is already). People should not do this either, even if it is also funny and relatively harmless. Also, suing people while participating in an attempted crime probably would not go well. =)
Have a gloriously wonderful day =)
Other options:
- A kind of mirror network (it only needs to keep sure that integrity can be checked, maybe with a public key)
- And while doing that why not also support compression (why not? only devs need to read it and they can run easily a decompression command), every bit saved would help.
S3 buckets in IAD with <5GB blobs can double-up as bit-torrent seeders.
I'd imagine, some tech IPFS/Filecoin/Sia might come in handy, too, but unsure of how healthy most of these web3 projects are right now.
There's also fosstorrents.com that help seed projects.
Sure it gained traction around Blockchain, crypto DeFI - but the storage technology is ELI5 a massive distributed storage.
I could hazard a guess that in terms of philosophy its closer to BitTorrent than Blockchain
Someone advertises the content to the network, then people looking for the content can find it. Usually the people who find it advertise that they now have it to the network too.
It has nothing to do with cryptocurrency but is commonly used as a great way to embed "larger" content into blockchains and other immutable stores. It works well for this because 1. The CID contains a cryptographic hash of the expected content 2. You can change where/how you store the actual content without updating the URL.
1. Try having it block subsequent requests for EasyList itself, just in case the frequent update requests are made with the prior blocklist in effect. (I accidentally did this before, in one of my own experimental blocklists, atop uBlock Origin.) Then the device vendor can fix their end.
2. If the blocklist language and client support it (I suspect they don't), you might safely replace or alter some Web pages, to add a message saying to disable EasyList in the client, or pressure the vendor, or similar. If this affects a lot of users, the meaning will also be spread in other languages to other users, even if not all of them understand any of the languages in the message. But be careful.
3. If you can't get a better message to the user, another option might be to block all requests, to prompt users to disable EasyList or vendor to fix the problem. But before doing this, you'll need to have verified that a combination of shoddy client/device software won't prevent users from using important functions of their devices for significant time. (Imagine this might be their only means of being connected online, and some shoddy client software pretty much prevents it from working, and the user is unable to access critical services.)
But before doing any of these desperate technical measures... First, I'd really try to reach people in the country who'll know what's going on, and who can reach and possibly pressure the vendor who's causing the problem. If tech industry people aren't able to help quick enough, reaching out to that government, directly or through your own country's diplomats/officials, might work. Communicating the risks of the desperate technical measures that you're trying to avoid (e.g., possibly breaking critical communications) could help people understand the urgency and importance of the situation.
Now you're thinking with portals
> When we encountered a similar problem last year, we found a simple solution: block the undesired traffic from these apps. Even so, we continue to serve about 100TB of “Access Denied” pages monthly!
Instead, you need to break the user experience so they complain to the developer of the app, thus impacting reputation.
It’s unfortunate that the browsers developers are unresponsive and this circumstance limits the available options to easy list.
I think something like this would be best served by moving to IPFS or Bittorrent. A magnet link could be provided and then browsers and plugins could use that to download the file. That way, you can distribute the load.
2. Send a cease and desist letter to the app creator.
3. If they don't respond, also send a C&D to Google demanding they cease distribution of the malware responsible for the DDoS.
Anyone can send a cease and desist -- it's just cautionary letter. You aren't obligated to follow through with the threatened legal action.
It doesn't have the force of law behind it, but it'll at least get their attention.
(IANAL)
So people writing apps using crappy coding are cratering Easylist basically through unintentional DDoS. It is the apps that suck. Full Stop.
I have noted over the last five to ten years the general retrograde usability of phone apps and full desktop apps in general. Bad UIs, inconsistent behavior, performance issues on stuff that, at least on the surface, appears to be trivial.
I don't understand what is causing the general slide in quality, but it is clearly visible, and it seems to me, untenable over the near term.
Is it the tools? Is it the app churn pressure? I really do not know, because I work in a very different part of the industry. We have our own issues there, but that has more to do with technology churn (particularly in wireless standards) than in the tools and platforms.
So what is up with the app world, in general, and the web app world in particular? I am all ears, because as I said, I don't work in that space.
To make the system more scalable: instead of directly serving the file, serve a bunch of URLs to mirrors plus a checksum. The client must pick one of them. You can randomize the URLs and maybe add some geo logic to it. Let people provide mirrors. An additional indiraction step like this can prove incredibly powerful for systems that need to scale massively.
Almost no dev would hotlink an asset that took that much longer to display, at least in critical/common paths. It would force consumers (devs/businesses) of the lists to provide a caching/mirroring solution of some kind for their users.
But on the bankend, the request would be designed just for updating the list cache. Handling 1-5 extra minutes per request, on a request that runs less than a few dozen times a day, to update the mirror/cache is trivial.
It was mentioned in this article that they are now serving up accessed denied, but the problem is one of just too many requests.
At this point, it's likely easier to just kill the domain all together and get a new one.
That being said the data is small so it probably isn't a big deal.
A better solution may be something like IPFS where with a rolling-checksum chunking algorithm you only need to download the changed parts of the list. You could even use IPNS to distribute updates for a fully decentralized setup.
You could try that instead of the "CDN" service
--
Alternatively, try the "cheap" CDN services like Bunny or Beluga, which have packages for high volume like 0.005c/gb
Cloudflare is not really selling a CDN, but all the "smart" services on top of it.
That's why you don't have as much control (like blocking IP/Geos without Enterprise), or run into issues for breaking their ToS.
You're off by a factor of 100.
https://www.belugacdn.com/cdn-pricing/
$5000/PB = $5/TB = 500c/TB = 0.5c/GB
> You pay 1¢ (or less!) for every Gigabyte of data accelerated over our cloud network.
doesn't seem clear cut as they bundle it in a subscription packages, but at this volume you'd surely need their enterprise package (10Tb+), which you'd expect to get 0.01 or less
bunny
> First 500TB $0.005 /GB
> From 1PB-2PB $0.002 /GB
That is several thousand $ a year.
The central HTTP server of easylist could then hand out 302s to fetch the actual file from one of the mirrors. Alternatively to 302s, a modern scriptable DNS server, that uses the mirror list, responds with different IPs, round-robin (or even better if it's geo-aware).
DNS TXT records could maybe be used to serve a digest of the file, do that mirrors can't modify it without that being detected.
Meanwhile perhaps some other CDN provider wants to create some goodwill if Cloudflare isn't willing?
I imagine some (many?) clients would handle it poorly, but it would then be cachable at least. It's not exactly easy to test though without a level of unknown damage to legitimate users.
The traffic issue is not just punted to DNS service. It's possible to return a cachable 127.0.0.1 response, and it's somewhat rare for DNS caches to be constantly powered up and down and reach out directly to authoritative DNS servers.
You then give them the finger, tell them to get their lists elsewhere, or pay for access due to their incompetency.
They will just have to do the needful.
What's the carbon footprint of this?
While still trying to allow valid traffic through.
The sent http body was blank, but I beleive we were still sending http head...
> If firewall access exists, just drop the offending incoming traffic entirely.
True, but the service we were using at the time didn't have a L3 firewall, and so we ended up moving out, after paying the bills in full, of course.
Is there a clean solution to this problem these days? Like some kind of adblocking router that resolves these addresses correctly but then routes packets destined for these services into a black hole so the requests eventually timeout? That would at least slow the repeat request floods down significantly.
And in this particular case even BitTorrent proper may not have helped because steady-state BT is distributed but if a client doesn't persist its state after a bootstrap - and lack of persistence is the issue here - it'd hit the bootstrap server every time. Granted, it'd only be about one UDP packet per client, much less traffic than what is easylist is seeing, but foolish code deployed at scale can still overload services provided on a budget.
proper solution would be to use DNS to forward all Indian traffic to one of the local VPCs.
The idea is to rate-limit users - let them pull blacklist.txt only once a day, so you serve the file and add requestor's IP to denylist on the firewall. So that any subsequent requests are blocked by firewall
+Cloudflare has Geofencing feature
Google playstore will remove them.
Trust in the process.
Pay for your own darn servers, everyone else has to and can't even use ads to support the costs due to you.
I am fine with one but not the other. Thus I block them.
If this browser is not maintained then it won’t update what it requests, but everyone else can.
Obviously this becomes very whack-a-mole but in this instance it might be an easy win.
Easylist should do something similar - requests from India should include a list with all the popular sites like Google, YouTube, aajtak.in, etc. When their browsers suddenly stop working the problem will be solved.
What is the reason for proxying through Cloudflare? Are there any bandwidth limits or performance issues when directly serving those files from GitHub?
---
> GitHub Pages sites have a soft bandwidth limit of 100 GB per month.
> In order to provide consistent quality of service for all GitHub Pages sites, rate limits may apply. These rate limits are not intended to interfere with legitimate uses of GitHub Pages. If your request triggers rate limiting, you will receive an appropriate response with an HTTP status code of 429, along with an informative HTML body.
> If your site exceeds these usage quotas, we may not be able to serve your site, or you may receive a polite email from GitHub Support suggesting strategies for reducing your site's impact on our servers, including putting a third-party content distribution network (CDN) in front of your site, making use of other GitHub features such as releases, or moving to a different hosting service that might better fit your needs.
---
'Legitimate' users of the list would clone/pull the repo to their own mirror?
EasyList updates frequently, many times each day, as the commits to that repo demonstrate.
I suspect the latter. I don't know how to make a repo public but limit web traffic to it. Do you?
You can serve that easily from a dozen dedicated servers for a low four figures amount. Or just get one server with a 10gbit/s connection if you don't care about being geographically close to everyone.
These numbers aren't large, but many CDN companies (and AWS) would like you to believe they are and you should be paying them insane amounts of money to serve that kind of traffic.
And given that this seems like a near existential threat drastic action seems warranted.
A 1000mbps dedicated server could serve it 70 MILLIONS times per day. Considering that most wouldn't be served (E-tags and whatnot), it can probably sustain a billion requests a day.
What am I missing?
A 1000Mbps server could only serve 10.8TB, and that's not even accounting for overhead/daily usage patterns/etc
https://web.archive.org/web/20220901000327if_/https://easyli...
If you're getting a 330KB file, maybe the server issues are causing the download to fail?
> The problem is that this browser has a very serious flaw. It tries to download filters updates on every startup, and on Android it may happen lots of times per day. It can even happen when the browser is running in the background
EasyList should be offered as a version-controlled copy you grab once, that then gets bundled with an app, rather than offering a download to be called from an app: https://easylist.to/easylist/easylist.txt (Currently down as of writing).
The only caveat is such a list needs to be updated, so then a version system should be implemented for EasyList and you periodically bundle the new version via app updates. It would save a lot of bandwidth doing this.
As a quick fix, there are many options for limiting per IP per timespan, e.g. fail2ban, you could configure it to punish bad apps without crippling functionality for others. Well, maybe crippling a little bit in some very special use cases, still better than it simply not working.
Podcasting 2.0 has been talking about podping as a solution because podcasting basically has the same problem with periodic polling of the RSS feed. Basically you subscribe and then receive notice there's been an update, THEN you go get it.
https://www.podcasthelpdesk.com/podping-and-other-stuff-with...
PFBlocker seems to default to once a day.
There is no approach you can describe that doesn't run afoul of the described badly-behaved browser app which willfully retrieves the entire file afresh at every init. If it can be downloaded, it will be downloaded directly by the badly-behaved mobile apps.
That shows the discrepancy between linear program execution model (on the classic desktop) and the quasi permanently running model of modern apps. Maybe the site should identify these rowdy browsers and give them a kick.
The normal easylist is way bigger and has lots of rules for ad blockers like uBlock Origin.
This is hilarious in an unbelievably terrible and tragic way. The scale is mind boggling.
I wonder which browser it is.
EasyList would need more than the 100mbs Scaleway offers for €2/month but Scaleway has more capable instances.
I would add a redirect to the makers of the browsers in question (so that the leechers got to deal with the traffic themselves; https://en.wikipedia.org/wiki/Inline_linking)
I might be ignorant of the scale of things but surely a 404 page can't weigh more than half a kilobyte right?
I think EasyList maintainers have every right to break stuff in this case. Also could easily turn the list into a HTML file as it's just as parsable as txt.
Alternative Lists @ https://filterlists.com/
If you baked something like IPFS or DHT torrent capability into the application and requested your blocklists that way you'd solve the single point of failure problem, but that's asking a whole lot from a shittily maintained and poorly configured browser fork.
Can that be edge cached with a future GMT expire date?
Alternately a 302 response to trusted local mirrors?
At a minimum, they should have either reduced the requests to something like a monthly download (not great, but far better than requesting a file every startup), or ideally, hosted and updated the file themselves, on their infrastructure. At least that would force them to look at their own hosting bills, instead of crippling a prominent and important contributor to ad-blocking software.
That's exactly what it is, and Cloudflare should be honest about that instead of coming up with an excuse as pathetic as ".txt files ain't web content".
They could do that anyway by putting up the "Checking your browser before accessing" page and redirecting to whatever file was being accessed. There's precisely zero need to outright modify the HTML files being served to inject such checks.
Tell me, how does CF put a 'Checking your browser' page in my `wget https://easylist.to/easylist/easylist.txt`
Bad things are already happening. Random HTML data is no less useful than no data at all.
> Tell me, how does CF put a 'Checking your browser' page in my `wget https://easylist.to/easylist/easylist.txt`
By serving the HTML instead (with JS that does whatever checks and then redirects to the intended resource), and if the client can't cope with that, then tough shit; the absurdity of such arbitrary evaluation of whether or not clients are sufficiently browsery to be worthy of consuming bandwidth aside, ain't "only browsers directly navigating to this will be able to cope with this and gain access" exactly the point of using such JS-based client checks in the first place?
A bit ugly, but works.
<sigh> and a lot of similar rebuttals to this one.
> In short, it's much harder to protect a text file against DDoS.
It's exactly the same difficulty. Bandwidth is bandwidth, bytes are bytes. The only suggestion I've seen where HTML/CSS/JS might be relevant is using JS for client-side probing, and if Cloudflare is indeed injecting arbitrary JS into HTML pages it serves then that's utterly horrifying and is a problem in and of itself.
> <sigh> and a lot of similar rebuttals to this one.
My point is that a lot of people don't seem to understand the 'why' of this issue and instead appeared to just jump to the conclusion that:
> It's exactly the same difficulty.
When that's not at all the case.
> using JS for client-side probing, and if Cloudflare is indeed injecting arbitrary JS into HTML pages it serves then that's utterly horrifying and is a problem in and of itself.
Well they are. From CF:
>> Cloudflare’s bot products include JavaScript detections via a lightweight, invisible code injection that honors Cloudflare’s strict privacy standards
But even before we consider that, if the request is for html then it's likely coming from a browser. If CF replaces that html with their own then the browser will likely run it allowing them to run all kinds of probes then run the redirect. The same is not true for a .txt file or an image.
We understand the "why" of the actual issue - i.e. the bandwidth consumption - just fine. It's Cloudflare's decision to micromanage data formats consuming that bandwidth (instead of just, you know, evaluating the bandwidth itself and being done with it) and using that as their basis for a ToS violation that's entirely whackadoodle.
> Well they are.
Then that is indeed absolutely horrifying and a problem in and of itself. Privacy concerns (of which there are a multitude) aside, this seems like a really great way to break all sorts of things, and I'd trust the claims of "lightweight", "invisible", and "strict privacy" about as far as I can throw them.
> If CF replaces that html with their own then the browser will likely run it allowing them to run all kinds of probes then run the redirect. The same is not true for a .txt file or an image.
The same is absolutely true for a .txt file or an image. Consider two flows:
1. You click on a link, your browser loads a text file.
2. You click on a link, your browser loads an HTML+JS file, the embedded JS redirects your browser, your browser loads a text file.
Same end result, same client-side probing opportunities, and without needing to rely on swinging JS into the end document like Patrick Bateman swinging an axe into his coworker's face while ranting about "Hip To Be Square".
Or better yet: just don't do this and only care about the raw bandwidth consumed, instead of actively making the World Wide Web a worse place with "clever" tricks like anally probing my browser to see if it's browsery enough for some arbitrary standard of browserness (for the sake of a "protection" that's almost certainly trivial to break with one of the umpteen headless browser solutions anyway).
It may be easy, and it may even be the only option, but it's a bad one that will need some thought from the maintainers I expect.
It's not ideal, but until the problem is fixed/better solutions are found, I think it's a good "first response".