Libgen Storage Decentralization on IPFS
freeread.org
freeread.org
The IPFS network itself cannot be taken down, but I wonder when will the publishers discover IPFS web gateways, and start filing DMCA takedown requests against them, and what are the gateways gonna do? Will Cloudflare implement a huge copyright blocklist or even terminate its service (or worse, send a list of IP addresses to Prentice Hall or Elsevier)? What about the semi-official ipfs.io gateway? And the smaller independent IPFS gateways (there are currently 10+ on the public web)? Will they be DMCAed out of existence?
Is that unusual?
Development of IPFS slowed to a crawl when most of Protocol Labs switched to working on Filecoin. It remains to be seen if either Filecoin or IPFS will ever be viable for their intended purpose.
(I say this, and yet I know it isn't true in practice. But why isn't it true?)
filecoin is also promoting in China. https://mp.weixin.qq.com/s/T0Qt2CEP7QsA6Zy9Vzvdkw
You can run the following in the console to see all your peers
for await (const peer of await node.swarm.peers()) { console.log(peer.addr.toString()) }
The console currently get's spammed pretty hard because they haven't implemented filtering out plain ws (non-wss) peers if you are connecting https. But it will use webrtc to find peers, and find all wss nodes in the network.The contact info for these servers appears to be retrieved via the DNS rather than the DHT. This isn't surprising because browser clients can't participate in the DHT, but the DHT is currently one of the biggest weak points of IPFS. IPFS's DHT is so chatty that it consumes over a megabit of throughput just to maintain an idle node. Lookups are excruciatingly slow and often fail completely.
$ ipfs dht findpeer 12D3KooWJxNHY6zE1KnzFdBJrTMCbhCVNtSLvjp7qR5Wsp49DnUC
/dnsaddr/ipfs.scalable.io
/dns4/ipfs.scalable.io/udp/4001/quic
/dns6/ipfs.scalable.io/tcp/443/wss
/dns6/ipfs.scalable.io/udp/4001/quic
/dns4/ipfs.scalable.io/tcp/443/wssI regret ever building a pinning platform on IPFS. The worst thing is that, when I expressed my frustration on Twitter about an IPFS release completely breaking IPNS updates, the founder called me a Karen.
The last time I tried it (maybe 9-12 months ago) it was a real resource hog (including choking the network) and really slow.
One could also imagine IPFS over Tor.
IPFS gateways (and Tor-to-Web) are reverse proxies, they are configured at the server-side to accept connections from the public Internet and route traffic to IPFS (and Tor), and they distribute content to the public Web. As far as I can see, this is a problem, someone can send DMCA notices, asking the gateway operators to take files down from the Web. The only protection I see is DMCA Section 512 (a.k.a Safe Harbor Provision), the operators are not liable for things solely hosted on IPFS, which is good, but gateways must comply with the takedown notices. So I think it's entirely possible for Cloudflare to put a huge blocklist on its IPFS gateway, and for smaller gateways that don't have the necessary resources to handle the requests, be DMCAed out of existence entirely. In the 2000s, RIAA launched campaigns against eD2k servers and BT trackers, including the use of honeypot servers to collect information. RIAA was not ultimately successful, but it created a major short-term disruption (eventually new servers would always appear in a different jurisdiction). A similar campaign against IPFS gateways sound possible if the publishers decided to launch an aggressive crackdown like the RIAA. The question is only a matter of whether the decision is made.
I'm not saying that IPFS requires the use of IPFS gateways, they don't, and in a sense, gateways are counterproductive for the decentralization goal. But they do currently provide an useful service for the Web, and I'm just speculating whether a major disruption is possible.
Well, then we might need something like IPFS-over-Tor, using Tors anonymity features to make it more safe for operators.
So you could validate it if you wanted to. Because the URL _is_ the hash, if you did validate it you could be very sure that it is correct.
If a webserver hosts a file, how can you be sure that file is correct? Sometimes the ship a sums file next to it. At best this can find corruption, but it would not find malicious files.
Maybe IPFS can introduce a feature to generate one-time URLs, e.g. `oneTimeURL = f(realURL);`.
Websites then can generate new URLs for example every hour or even for each client request. Copyright bots will file a DMCA request to a URL which only they have.
It kind of worked for MegaUpload. Their founder got rich and a version of the company still exists. However, the founder is also gradually losing a long court battle and may find himself extradited to the US where he will face significant jail time.
It's a politico-technological arms race. Governments make laws. People don't agree with them so they create technology that subverts the government. The state must then become even more tyrannical just to maintain the same level of power. We'll end up with either an uncontrollable population or a totalitarian government.
I don’t know, do you think nobody has noticed LibGen by now? Do you think the reason they’re still running is because of some jurisdiction issue? Even the lawyers would be wrong on that one.
I'm not sure about the implication of your words. But, clearly, Elsevier has already tried to take legal actions against LibGen previously, and once in a while, LibGen's domain names are still being blocked and routinely replaced, just like Sci-Hub. I'd still say the reason it's still running is a jurisdiction issue, or let's say it's a geopolitical issue, it's using Russia to shield itself away from US influences.
But an IPFS gateway hosted in the United States by Cloudflare does not enjoy these legal and political benefits.
It's a bit of tricky problem, because people may not like unknown, random files on their pc, and probably, because such algorithm is distributed and will need to resist attacks.
But i think that the payoff is big: what Bitorrent did to popular files, such an algorithm could do to rare files.
IIRC eDonkey prioritised rare files, although I don't know if this was a property of the network or just a common choice in clients.
We also found it necessary to build a kind of "overlay" to the p2p network (how each peer selects its peers) because the default algorithm produced too sparse a network. The probability of one node ever being connected to another node that holds content it would like to sync approaches zero. The topology for the overlay network is also fetched from said blockchain.
It’s good stuff nonetheless.
[0] https://github.com/ipfs/community/blob/master/code-of-conduc...
The CoC you've linked covers the IPFS Community and other Protocol Labs initiatives, it doesn't cover usage of IPFS itself. freeread.org is not bound by that CoC.
I guess TFA is Teach for America? Doesn't really matter, but anything that is linking to IPFS Desktop could easily link to their own distribution of the client as it's all open source.
And while IPFS Desktop could start blocking content, so could torrent clients, not sure what's different really. They both are just clients for a protocol anyone can write new clients for and both have the risk of adding blocklists and having people abandon the client, so the same risks exists for both of them.
> so could torrent clients, not sure what's different really
This is precisely why the developers' attitudes do matter. The difference is that the BitTorrent community has a well-established permissive attitude toward copyright violations. This is what gives me a high degree of trust that my next distro update won't bring in any new content blocks. I don't have that high degree of trust in the IPFS developers, given the explicit statements made in favor of upholding existing copyright law.
> anything that is linking to IPFS Desktop could easily link to their own distribution of the client as it's all open source
Sure, but getting people to migrate isn't easy. Look at how many people are still downloading OpenOffice vs LibreOffice, for instance: http://www.openoffice.org/stats/downloads.html
What does that have to do with the network? Again, the network nor the protocol currently has nothing in place that describes a "attitude" of being against copyright infringement.
I remember that we (I'm a ex-Protocol Labs employee) used to have plenty of discussions about adding allow/blocklists to the protocol that people could opt-in to, but don't think that was never added (yet?). If it was, I'd understand where you're coming from.
Remember youtube-dl got taken down because it's README file referenced copyrighted material.
It doesn't matter if a tool _could_ be used distribiting copyright material. The tool (and communities) cannot promote anything related to that use case at all.
The same is true with bittorrent. Bittorrent _could_ be used to download copyrighted music and movies, but you won't find a single reference to anything like that use case on their web page because that is illegal: https://www.bittorrent.com/
In fact the bittorrent terms of service says you may not use the tool for any illegal purposes: https://www.bittorrent.com/legal/terms-of-use/
Everyone has wishes, but wishes aren’t enforceable. It’s the equivalent of “Stop! Or I’ll say stop again!”
> they've
The IPFS Forum - Run by Protocol Labs
> decided to blacklist
Banned discussions around breaking the law on a public internet property owned by a US company.
None of this concerns the protocol itself, which you are free to use for whatever you want, just like HTTP.
Even if so, what can they (the IPFS devs) even do about it?
Also, Freenet actively seeks out data from the network to store locally, to give plausible deniability to users (there's no way to tell if something is locally cached because the user requested it, or whether the client fetched it automatically). IIRC this can be disabled by limiting oneself to a whitelist of peers (preventing network chatter that might reveal Freenet's presence). IPFS will only cache what's requested, so it's more predictable and useful for building infrastructure (e.g. hosting assets for a normal Web site); and presumably leaner on data requirements too.
IPFS isn't really competing with Freenet IMHO; it's competing with HTTP and BitTorrent.
It has become a lot faster recently, try again :) When I look at its statistics during active usage it can easily show 1 MiB/s traffic.
You still can't expect HTML to load in sub-second delays, for large multi-MiB sites it may take a low single-digit amount of seconds, but there is an inherent cost to anonymity which probably induces a boundary on how fast things can be:
To achieve anonymity, data needs to be redirected across multiple people so the sender and recipient can't determine who each other of them is.
Redirecting stuff across a longer path than necessary slows things down.
>
> Redirecting stuff across a longer path than necessary slows things down.
That's skipping past a lot of details. There are still choices there. Freenet traded speed, efficiency and scalsbility for a specific security model. There are anonoyminity systems (e.g. tor) that are much more efficient but have different trade-offs.
I'm not sure why IPFS is more popular than Freenet, but the difference mentioned above certainly makes me more likely to run a IPFS node than Freenet.
Naive question: is there such a thing as Bittorrent over I2P or otherwise anonymized decentralized file sharing?
Yes. I2P is designed with P2P in mind. The default Java I2P client already comes with a bundled BitTorrent client, there is also a BitTorrent client called "Vuze" and is used by some non-anonymous users to help seeding a torrent across the clearnet and I2P simultaneously. Because of the nature of I2P, don't expect great speed. No seeder is "usual", 10 KiB/s is "normal" and 50 KiB/s with only 3 or 5 seeders is "good", overall, it feels as if it's still the early 2000s, but when you get used to it and adopt a "download and forget, come back in 3 days" mentality, it's not actually that bad. The fastest download I've ever seen was the movie Snowden, 20+ seeders and over 200 KiB/s+...
IPFS solved the data availability problem because you can incentivize storage miners to make your data available. This is something new, IMO.
Accessing it anonymously is possible with an extra protocol layer.
> The InterPlanetary File System (IPFS) is a protocol and peer-to-peer network for storing and sharing data in a distributed file system. IPFS uses content-addressing to uniquely identify each file in a global namespace connecting all computing devices.
And I was going to try this one of my embedded computers :(
As for piracy if the places I see libgen and it's kind referenced are any indication then this isn't about just any readers but about students, professors etc. who want to access scientific publications.
Torrenting copyrighted materials (especially if you are distributing) can be quite expensive in some jurisdictions and I do not believe that IPFS is treated different than regular BitTorrent.
It works well for file sharing as well. Don't know about copyrighted material though, have only shared my own files.
Why? Your ISP charges for BitTorrent traffic?
Edit: I didn't downvote you if it looks like that :)
And you have to pay the fines? That's nuts.
In the US, you are pretty much as likely to get fined for torrenting as you are to get arrested for jaywalking (although without the racial disparity).
Yes, conveniently there's a second group of lawyers that focuses on getting people out of these fines...for a fee that's not so much lower. Unfortunately it's a bit of a business model in some countries.
This is not "very" common - the ISPs don't really ban you, they just threaten you. I no longer do this, but I extensively torrented over multiple home ISPs and they would send me warnings but they would never actually do anything.
Source: wikileaks
Though I'm not sure how big of an issue this is; sci-hub is using https.
Also, Let's Encrypt doesn't even revoke certificates for malware and phishing sites. [1]
Maybe they're doing this for better caching performance? It makes SSL flood attacks impossible, but is that really worth it?
[1] https://community.letsencrypt.org/t/how-to-report-abuse/4110...
comics is not a public dataset -- it's privately run and managed, so there are no mirrors for it. I guess someone could scrape it.
It's hard to believe, but it's certainly happened/happening -- plenty of democratic countries are filtering the internet, like Turkey.
VirusTotal includes some very poor quality Chinese AVs that produce positives for nearly every file. PDFs are pretty damn safe to access today -- the days of exploding macro exploits are about a decade behind us.
I'm not surprised they're now flagging what is basically text and images as malware.