Cloudflare disables access to ‘pirated’ content on its IPFS gateway
torrentfreak.com
torrentfreak.com
Cloudflare is fine with Chromium from the same residential IP address, however, but I don't want to use Chromium for those Web sites.
I also don't want to have to try to debug the obnoxious behavior of some third-party company that perhaps doesn't care whether it's blocking legit users from its customers.
The Kafkaesque Cloudflare "prove you're a human" infinite loop is hopefully not a foreshadowing of an imminent Internet dystopia, with Cloudflare the vanguard of it.
Totally frustrating that Cloudflare is basically the Internet's biggest gatekeeper with a mysterious black box behind it that punishes legitimate users.
Every site you put behind Cloudflare contributes to the future where they have the singular ability to decide what networks can connect to others.
Except for the fact that you can disable Cloudflare on your site anytime you choose.
You might be interested in a plugin to toggle the fingerprint protection: https://addons.mozilla.org/de/firefox/addon/toggle-resist-fi...
> The Kafkaesque Cloudflare "prove you're a human" infinite loop is hopefully not a foreshadowing of an imminent Internet dystopia, with Cloudflare the vanguard of it.
The "prove you're a human" glitch, should be a really obvious fix, that lots of users have continuously complained about for a while. You would think that once proven human, that means access to the site. Somehow, Cloudflare has shown no interest in fixing it. In fact, they appear to have more interest in destroying user privacy and protections, using access to a site as the carrot. A lot of the webmasters seem to be getting overzealous or don't have a clear understanding of how they affect users. Cloudflare appears to not be helping, but rather pushing tools and settings to promote sales, so we get user nightmares like "prove you're a human" infinity glitches.
> Apparently, that’s not the case for IPFS.
But it’s also not the same situation at all.
With DNS, blocking would mean making the whole site unavailable.
With their IPFS gateway they are just filtering individual items.
For IPFS vs CDN, I'm guessing relevant copyright laws give a service provider an out when it's a well-defined user/customer doing the copyright-related redistribution; no such agreement exists for content on IPFS.
EDIT: Trying to read the relevant DMCA sections and precedent makes me happy I'm not a lawyer. Cloudflare was previously sued for copyright infringement (Mon Cheri Bridals, LLC v. Cloudflare, Inc.) and found not guilty. In Cloudflare's own blog about the decision, they said the suit was meritless for several reasons including "our services are not even necessary for the content’s availability online." My guess is there is sufficient ambiguity over whether operating a inter-protocol gateway puts them at increased legal liability.
Unless I'm missing some magical internet tomfoolery, that's not the case when you use their DNS proxy service, right?
That's their CDN with a twist.
For DNS specifically, it is analogous to saying a phone book shouldn't be liable for listing the phone number of someone who sells pirated DVDs.
Guess I could just try deploying tornado-cash..
The site is still available, just not through its DNS entry.
> With their IPFS gateway they are just filtering individual items.
They can filter as broadly or as targeted as they want.
A public list of what their gateway will not retrieve for you would fuel the Streisand effect...
hxxps://libgen.is/dbdumps/
That sounds more like a limitation of your bittorrent implementation. You think seeders have dozens of identical copies of each release hanging around on their drives just to be able to share it on separate trackers through separate torrents? :p
Or is there some new feature in IPFS that's an actual differentiator here? It's been a while since I was fully up to date on protocol development.
> BitTorrent v2 not only uses a hash tree, but it forms a hash tree for every file in the torrent. This has a few advantages:
> Identical files will always have the same hash and can more easily be moved from one torrent to another (when creating torrents) without having to re-hash anything. Files that are identical can also more easily be identified across different swarms, since their root hash only depends on the content of the file.
So I guess the question stands, how is IPFS supposedly the preferred protocol over Bittorrent (v2)?
You could make the experience similar though if you really wanted. So in theory, no difference. In practice, ipfs solved this years ago, has a nicer interface to it and put updateable naming on top for convenience.
Isn’t that the purpose of BitTorrent’s DHT?
There is actually a conference of ipfs devs and users in Brussels going on (or held not long ago). https://2023.ipfs-thing.io/
I have no idea of who or what pinned the ipfs content. I really don't gather those statistics and have never tried to do so. The only "metric" which grabs my attention is that the content is widely distributed and that its reliably accessible via more than one way (not depend on a single point of failure) over large-scale electronic networks/protocols (HTTPS, DNS, IPFS, ETH, ENS..).
Insofar as I understand IPFS, the most natural use case I've come across would be serving Nix packages.
That is, my understanding of IPFS is "protocols related to distributing content based on its content". Nix stores all its packages in a nix store, where its address there is a hash of the inputs; but, there are experiments for addressing Nix packages by the package's contents.
https://www.tweag.io/blog/2020-09-10-nix-cas/
On the Nix side of things: a cache-miss from a binary cache would just mean that a package would need to be built from source.
Speaking of which, is Freenet still alive?
I saw the writing on the wall when my ISP was forced to block TPB and KAT.
A seedbox is a temporary protection. In the long term it is still a very real possibility that the customers of seedbox companies can get in trouble for their acts of piracy.
Several VPN companies who claimed to keep no logs, were actually keeping logs. https://www.pcmag.com/news/7-vpn-services-found-recording-us...
So too it will turn out that some seedbox companies will have been keeping records about their customers and their activities. For those customers, it will not be a happy day.
It may be a stupid law, it may be an unjust law, but it's still a law and people with guns who enforce the stupid laws might show up some day.
Only individuals can decide whether or not the risk is worth it.
Unfortunately, you need to anonymize your traffic (via a VPN service, which costs a fee) if you don't want to be harassed.
> Free from censorship. Free from privacy violations.
I2P is neither “for” nor “not for” piracy. I2P is for protecting communications and data from any kind of censorship.
Filtering copyright violating content is a type of censorship. To some it is justified. To others it is not.
To a network protocol like I2P the point is that no kind of censorship should be possible. Regardless of reason.
Why would you expect customers to migrate someplace else at all if there's no alternative?
Of course, having DDoS protection on the machine being flooded is a non-starter.
DNS would have to be highly available as well.
And it all needs to be maintainable by someone who's main job is not maintaining infrastructure, and doesn't have the expertise. Assume hiring someone to do so is not in the budget.
DDoS management is harder. You just kind of need to assess your risk and take appropriate steps given your risk. If you're likely to attract real, determined, attacks, you need a good solution.
If you're going to just get bored idiots that control a lot of bandwidth, but don't have a real beef, accepting that you'll get null routed for some time when you get attacked is really the least hassle option. Get as big of an incomming connection as you can justify, make sure you can discard bogus packets at line speed, take steps so that you don't amplify repsonses, and cross your fingers.
Be aware that moving to new hosting while under attack can be difficult (most hosts do not like customers that attract DDoS, especially new customers), so if you do move, communicate the current situation to the new host beforehand.
If you have important services (APIs or what not), offer them on hostnames other than www or your apex domains. Idiots flooding random people for lols really like to hit www, and ignore your other hostnames.
Free basic services arn't a loss leader, they're paid for by the enterprise contracts.
The loss leaders are always paid for elsewhere, or at least that's the plan.
Not saying you don't. Saying your experience is not everyone's experience here.
I can't afford to serve the 98% that is bots. That's about 10 search requests per second, most of them are search queries aimed at poisoning the query log of my search engine, which doesn't exist. But I think they're just scraping for opensearchdescription definitions and spamming every endpoint they can find in the hopes it's wired up to Google or Bing.
Cloudflare are the only ones that seem to offer some sort of (affordable) mitigation against this. They're hiding behind a botnet so rate limiting and ASN/IP blocking does all of bupkis.
The other option is to shut down my website. I believe that world is a worse world than one where stuff gets routed through Cloudflare, although I'm fully aware that this is far from ideal.
What I do to help those who want to avoid cloudflare is offer API access, which means I give you a token so I can rate limit you. Also means less anonymous of course, but it's at least free from men in the middle.
What benefit do those bots get from "poisoning the query log"? And what benefit do they get from "spamming every endpoint ... wired up to Google or Bing"?
The fact of the matter is they're basically spamming very specific search queries and not really looking at the results.
But most larger search engines store and analyze search queries to for example produce typeahead suggestions or error corrections. You can conceivably poison this data to manipulate how a search engine operates, perhaps directing users to your website or affecting ad auctions.
I think the benefit of not directly spamming Google or Bing is that they're harder to identify.
But this is a publicly available Internet search engine we're talking about. There's slightly more processing involved in processing a search query than loading a static file off a filesystem.
So what's happening here could be one of two things:
Link injection, where by spamming an endpoint with a URL they hope the URL will end up indexed and displayed somewhere on the site, for a search engine to pick up later. Often with some keywords in the link text, etc. The idea is to boost pagerank.
The other is even sillier: they try search for their own sites a tonne on third party search things that probably query a "real" search engine, to try artificially boost the supposed popularity/keyword associations of their sites, boosting pagerank.
A lot of blackhat SEO stuff us bunkum and snake oil.
Every so often they figure out a working method to exploit search engines to boost pagerank, and then it gets fixed. But they never stop trying even the most obsolete methods as the automated scripts are what they have available.
The only place that filter exists is where people actually communicate openly about their strategies.
As described on the Security Now podcast, this is now part of "internet background radiation"
I can't host a website for my home internet without the risk of getting black holed.
Can't ask the website for my home internet because Comcast won't let me.
Providers will charge you out the ass for bandwidth if someone ddoses you.
The internet just isn't a sustainable place for self run websites now,
I wish residential providers would support a second level of quota management. Assign 90% of my bandwidth to this website, then take it down for the month without completely killing my internet.
So this isn't insurmountable. You don't need a centralized provider for protection from decentralized abuse.
Having my website instant-killable "for fun" by any bored script kiddy living in a country thousands of km away from me is not a particularly interesting proposition.
There's also tons of sites that get the "HN hug of death" when they're posted here (or Reddit, or wherever else that generates a surge in traffic). Tragic to miss out on all those potential hits and easy to prevent with Cloudflare free.
You can dislike Cloudflare policies, but they do provide a valuable service and often completely free of charge. That's exactly why they have such a large slice of the internet under their domain. People don't just opt in for no reason.
Sure. But I was originally responding to a comment that claimed the key reason to use Cloudflare was DDoS protection. If the traffic is real, it's not too unreasonable to pay for it.
Using any service comes with costs and benefits. Any reason you have to use a service irrespective of their behavior is just a blind spot.
With TLS it’s arbitrarily easy to overwhelm a target site that’s hosted on the smallest instance types at $provider due to the compute requirements of terminating TLS. Thankfully Cloudflare + LetsEncrypt makes it economical to host a personal site with good security without arbitrarily and suddenly high bills, constantly high bills, or my site arbitrarily disappearing off the web.
Bad actors poisoned the well (thanks China), and now here we are. I, for one, appreciate greatly what Cloudflare has done for the indieweb.
https://www.bufferbloat.net/projects/codel/wiki/CakeRecipes/
Yes. I could.
This is one of several possible solutions available to me as an alternative. All alternative solutions to this problem either:
1. Are more expensive to implement, for various values of "more expensive".
2. Are more complicated to implement, for various values of "more complicated".
3. Are less resilient to the problem space, for various values of "less resilient".
In some cases, they are all three. The proposal you made happens to be all three.
If I were to do things all over again from scratch /today/, I might do things differently, simply because I'm more well off and have a more stable physical location than when I set up the last iteration of my personal site (which is still running), but none of the possibilities I can think of are strictly "better" than what I have already done, leveraging Cloudflare.
What would need an experiment to prove that? You would just have the proverbial receipts
Or do you have proof that they’re actually ddos in people?
Step 2: Enable cloudflare, paid tier, hang around with it turned on for a month or so
Step 3: Simulate some usage traffic
Step 4: Cancel your cloudflare contract
Step 5: Watch your unknown site get ddosed
Coincidence? Suuuuuure
Cloudflare itself doesnt need to ddos you, plenty of ddos for hire outfits out there that will happily do it.
As to who hires them, well that's between them and their customers. But there are some things that are not coincidences.
I know that Meta is doing its level best to wipe LLaMA's weights from the face of the Internet, but I also suspect (and cannot prove) that various organs of the US state (CIA, NSA, FBI) are also trying.
If they're not, I'm even more afraid, because one of their most important jobs is being alarmed so I don't have to be.
I'd wager a few dollars that Cloudflare had a visit.
And, as an aside, I'd say it's a matter of when, not if, a GPT-4-scale model is leaked, possibly GPT-4 itself. If not to the open Internet, then to China.
OpenAI is (for now) staffed by humans, and humans are vulnerable to seduction, spearphishing and spycraft.
The idea that Meta is frantically running around trying to "wipe" these weights is ridiculous. They're either doing this to disclaim liability in case a new law imposes it retroactively, or they are trying to create uncertainty around its use to dissuade serious US competitors from using it for their own products.
> but I also suspect (and cannot prove) that various organs of the US state (CIA, NSA, FBI) are also trying.
I would be very surprised if they were given that it would only handicap the US while leaving other countries to advance unabated. Again, the only real thing facebook did here is show that we can have smaller models that are competitive to larger ones. The information to build and train these models is public. The code to do this is public. The capability to this so is public.
The cat is out of the bag and humanity is better off for it. What humanity needs to do now is figure out how best to adapt.
Do homework. SMH.
IPFS is primarily used to store infringing content, yet few users directly interacted with IPFS. Instead, they usually rely on Cloudflare basically serving as a free proxy.
It's like the early days of mega or something, just with a layer of indirection that serves as plausible deniability.
Why Cloudflare bothers to provide this presumably costly service at all is confusing to me. I guess maybe they drank the crypto kool-aid and thought that it might some day be profitable?
IPFS is crypto adjacent and got plenty of hype because of it.
That doesn’t mean I love the way that Facebook influences the world – quite the opposite, but that means I have even less tolerance for people hawking get rich quick scams as the solution since they’re distracting attention and resources away from things which can possibly work. It’s similar to how cryptocurrency salespeople talk about “the unbanked” as if the solution to that problem is to make a bunch of rich people in the developed world even richer in the hopes that they’ll share some of that largesse.