Putting Z-Library on IPFS
annas-blog.org
annas-blog.org
Z-Library - "the desktop app", with built in Tor (for seeders safety), IPFS (for p2p distribution), IPNS (to download updated indexes), with local search engine (no SPOF & convenience), optional at-rest file encryption, and some random pining algorithm to let users donate 1-10GB of local disk space to host random chunks of the library.
It needs to be dead easy to let anyone use and contribute.
The main demographic is people who are probably already technically inclined... but there are lots of supporters of intellectual freedom who are not the main demographic of Z-Library.
This is a provably unsolvable problem, but it's okay: nobody actually has this problem but the worst social outcasts. In some places that's itself a problem: what "The worst social outcasts" is varies from place to place, but running IPFS is enough proof for a death sentence in those places.
As far as I know, there's no such statistical anonymity in it.
I think the parent comment is asserting that even if you run IPFS over Tor or something, or even if they try and make IPFS more anonymous in the future, it will always still be vulnerable to the "cut off your Internet and see which node in the network goes down" attack.
Perhaps we could narrow down the purpose of IPFS to storing a small amount of data in a database that is impossible to delete and difficult to censor?
To my knowledge, IPFS isn't really private, in that both the nodes hosting content can be easily known, and the users requesting content can be monitored. This is bad news for something law enforcement has already taken a serious interest in.
IPFS also requires "pinning", which means that unless other people decide to dedicate a few TB to this out of their own initiative, what we have currently is a single machine providing data through an obscure mechanism. If this machine is taken down, the content goes with it.
The amount of people that have 31 TB worth of spare storage, care about this particular issue, and are willing to get into legal trouble for it (or at least anger their ISP/host) is probably not terribly large. The work could be split up, but then there needs to be some sort of coordination to somehow divide up hosting the archive among a group of volunteers.
You can access the data through a VPN.
If necessary, the hosts can also encrypt the filenames and data so that, until law enforcement gets the encryption key, they can't know who accesses what (public key and other necessary info would be communicated through Signal). Rotate the filenames so so when one is discovered, past requests can't be tracked. Maybe there is a way to slightly break the protocol to further hide the requests.
> IPFS also requires "pinning", which means that unless other people decide to dedicate a few TB to this out of their own initiative, what we have currently is a single machine providing data through an obscure mechanism. If this machine is taken down, the content goes with it.
Do you have to explicitly choose what data to pin? If so then this is an issue. If not, and you just pin random chunks, then if we normalize people using IPFS and distributing legal data this will be solved. If we normalize it enough, there will be too many people hosting and using IPFS for law enforcement to reasonably take down. Or we could just have enough activists that are willing to risk being fined or arrested.
---
That being said though, I'm still not convinced on IPFS because it seems like it cannot handle much and is excessively inefficient (case in point: this article). The authors of IPFS should release a new protocol which addresses issues like the article's, hopefully before too much adoption.
Okay, and then law enforcement asks the VPN. Yeah, it improves matters some, but we're talking about a huge book archive here. People aren't going to maintain OpSec
> If necessary, the hosts can also encrypt the filenames and data so that, until law enforcement gets the encryption key, they can't know who accesses what (public key and other necessary info would be communicated through Signal).
That's a plan suitable for some sort terrorist organization maybe, but exactly how is that going to work for an archive of millions of books that are intended to be served to the general public? What's the key distribution mechanism? How do you distribute keys to everyone but the cops?
> Do you have to explicitly pin the data? If so then this is an issue. If not, and you just pin random chunks, then if we normalize people using IPFS this will be solved. If we normalize it enough, there will be too many people hosting IPFS for law enforcement to reasonably take down.
IPFS isn't Freenet. My understanding is that it's a content-addressable, multi-source system. Meaning the main different thing from plain HTTP is that stuff is named by hash, and that if there's a dozen people serving a given file, then the system can spread the load among them, or tolerate some of them going offline. You ask for hash X, the system figures out where to get it.
Unless people make the intentional choice to mirror content, then it's not very different from serving stuff over HTTP, only with a worse user experience.
What you suggest sounds more like Freenet, but I doubt that it'd work great even there. Freenet does the "store random chunks" sort of thing, but this means that it's extremely inefficient, and easily loses data. Freenet was made for plausible deniability, so any storage is probabilistic, and data is replicated as it moves through the network and eventually lost if nodes go offline or it just falls out of storage due to the lack of interest. Storing 31TB would require a lot of nodes dedicating a lot of storage, and a lot of interest in accessing all of that data on a regular basis.
> If we normalize it enough, there will be too many people hosting IPFS for law enforcement to reasonably take down. Or if we just get enough activists that are willing to risk being fined or arrested.
That's not a great plan for something that already got people into legal trouble
I’m only saying this because I suspect that otherwise if people fail to change your mind you’ll walk away thinking you were vindicated by their silence, when it’s just as likely that they’ve followed the old adage: “Never Wrestle with a Pig. You Both Get Dirty and the Pig Likes It”.
That's exactly what the person suggesting using VPN and encrypting file names is doing.
You're making it seem like this is such an impossible task, but libgen already uses ipfs as a mirror..
It is better than nothing in the circumstances, but be aware of the species of creature it is.
Seems like over the last 5 years they haven’t really done anything but add crypto buzzwords to the project site.
As other commenters pointed out, the biggest real difference is that IPFS objects can cross-reference each other by hash, which allows you to do partial updates.
You can think of each CID as a mangent link, and each CID can point to more CIDs.
Where as in a torrent the peer tracking is done at the whole torrent level, instead of the chunks inside the torrent.
Torrent V2[0], which has poor adoption thus far, should also allow for a similar ability. To my knowledge this is not being taken advantage of yet.
An extremely interesting problem, fingerprinting generic binary - reducing kilobytes and megabytes into a handful of bytes, and avoiding collision.
Much simpler on large data.
Anyways, if 6 million ebooks takes up 31TB, I wonder how much space the equivalent amount of audiobooks would use? And how much space a decent library of movies has? I'm kinda thinking that guy on reddit who somehow got a hold of a used Netflix edge server was on to something.
id say change the audio format and cut the bit-rate roughly in half.
but welcome to data hoarding, start with a 8tb external wd dirve!
1. https://www.audible.com/pd/Sherlock-Holmes-Audiobook/B06WLMW...
Quick counterpoint to this type of revolutionary rhetoric - writing a book (a good one) takes a lot of time, patience, and hard work. Making it available for free, against the author's wishes, while very easy in the digital age and obviously great for readers, is robbing from the author (if they're still alive). It devalues the craft of writing and makes it even more un-viable as a profession.
This may be inevitable, and it has mostly happened with music, journalism, etc., but by putting energy into this kind of project you're only furthering its death along, and perhaps furthering the death of society and culture a little by taking away a lot of the incentive to write anything of value.
Pirate away if you want, but thinking this is some kind of virtuous revolutionary act is narcissistic BS.
https://www.ibiblio.org/ebooks/Lessig/Free_Culture/Free%20Cu...
I wonder if the answer to these questions also resolves the unfounded idea that writers will stop writing when capitalism ceases to exist.
Basically you're saying authors should be forced to work for free, with day jobs to pay the bills, and writing on nights and weekends, rather than having their craft be a viable profession.
Fine, and it's probably where things are going, but kinda sad in my opinion. And the hypocrisy of book stealers thinking they're some kind of Che Guevara is pretty silly.
Where did GP make that claim? Closest I found was that copyright violation is "taking away a lot of the incentive to write anything of value." That may or may not be true but they certainly did not say that "writers will stop writing when capitalism ceases to exist."
What on earth are you talking about? Some of the best artworks - in all categories mentioned, and beyond - come with a flippancy towards marketability, aka the avant-garde. The detachment of capital from expression will only prompt more earnest expression.
Yes business can cheapen and corrupt, and produce commercialized crap (which apparently people like and will pay for so it's serving some need). But there are plenty of artists out there with integrity who also make a living at it.
Is this an “Artists should come from rich families“ or “Be like Picasso and Warhol. Attend the right parties with the right people and you can attract the right patrons too.”?
The only people who can be flippant with regards to marketability are those who have family money, an extremely supportive spouse, a patron or are willing to live in poverty.
Art has always been pursued under these exact circumstances, and literature was no exception historically. Mass-marketable art is very much the exception, not the rule. And often it's even the least interesting kind, because it's the most predictable in its features - more of a skilled craft than art in a narrowly creative sense.
In this aside, you just described most artists. Picasso and Warhol are the two biggest pop stars in art history, and the worst possible examples of what constitutes the average artist.
Can you explain a bit more what you mean? How does this justify copyright violation? Unless you ask the author, how do you know they would let you read the book for free? Surely they get paid more when people buy the book, right? IIRC most authors recieve royalties.
Royalties often come out to pannies per book. Even an advance of a few thousand dollars ends up taking tens of thousands of unit sales to cover the author's advance and have them see any direct royalty payments.
Publishers also pull all sorts of shady accounting to stiff authors on royalties. They will charge them for editing or cover art against their royalties. Worst is when they charge the author if a retailer returns unsold copies.
That's all besides the fact Z-Library is little different than any other library. A library will lend out a book a large number of times and only pay for it once. Most of the readers were never ever going to buy the book and only read it because it was available at a library.
Very few authors ever see any money from royalties and fewer still manage to live off royalties.
>fact Z-Library is little different than any other library.
The difference is that copyright does not cover benefit from a work but reproduction. Libraries would be illegal too if they involved reproduction of the work.
In any case, this is all still far from the "it was sad day for the free flow of information, knowledge, and culture" argument given by the OP. Most authors really don't want their works reporduced at no cost, and this idea is basicaly a "what's yours is mine" mentality.
Publishers make money off selling books. Old books do not sell (obviously there's outliers). Publishers offer advances to get new books to sell. If they stop offering advances they won't get new books.
For most authors royalties are not a thing and never will be. Worrying about royalties on their behalf is pointless.
None of that changes the fact that pirates are no different from a non-customer. Equivocating about "reproduction" is just ridiculous. A library loaning a book out to a hundred non-customers is no different than a hundred people pirating a copy of a book or borrowing a copy from a friend. They're all non-customers.
Treating pirates as some special class of non-customers is ridiculous.
>For most authors royalties are not a thing and never will be.
Can you provide some numbers to back up these claims?
Also, I'm not clear on where you addressed my point that this would "long-term result in publishers offering less advance money." If it's that "if they stop offering advances they won't get new books" how do you know that they are already offeringing them exactly the minimum such that if they give them smaller advances they won't write more books? Surely it would depend on how much they think they will make from the sales? And what about self published books? Do you admit that those should not be pirated?
>None of that changes the fact that pirates are no different from a non-customer.
Do you mean on the avergae or in totality? Because there certainly are people who would pay but instead pirate.
>Equivocating about "reproduction" is just ridiculous.
That is what the law of copyright is.
>A library loaning a book out to a hundred non-customers is no different than a hundred people pirating a copy of a book or borrowing a copy from a friend.
Authors do not have conrtol of what buyers do with their books beyond the fact that they cannot copy it. That is the condition of the sale. It is usualy expressed something like:
>Copyright © [year] by [author]
>All rights reserved. No part of this publication may be reproduced, distributed, or transmitted in any form or by any means, including photocopying, recording, or other electronic or mechanical methods, without the prior written permission of the publisher, except in the case of brief quotations embodied in critical reviews and certain other noncommercial uses permitted by copyright law.
Basicaly I'm making two points here. One that authors do net lose money from piracy, and second is that even if not, they expressly told all the buyers that the condition of the sale is that they cannot reproduce their works.
As for real evidence for my claim, I point you to the Author's guild[0] which fought against google[1] providing for free copyrighted works even if it contained a link to their store. And it in general defends author's copyright. It "has counted among its board members notable authors of fiction, nonfiction, and poetry, including numerous winners of the Nobel and Pulitzer Prizes and National Book Awards. It has over 9,000 members."
[0]https://en.wikipedia.org/wiki/Authors_Guild [1]https://en.wikipedia.org/wiki/Authors_Guild_v._Google
It's not a zero-sum game, where either the author wins or the middleman wins. It's complicated.
Your mistake is putting the author first in this list.
For much of human civilization there has been an inherent social contract of sorts around communication.
I can take what you tell me, mix it with other things, and retransmit it. While I love books, bits and bytes are not books. There is no physical media. Now, we are in an era of inherently zero cost redistribution, and the vast and giant fraud perpetrated on us by the cabal of distributors, publishers, authors, publishers and agents finally is seeing pushback.
The vast amount of exploitation of public research, and non-value adding commercialization inherent in some of the scientific publishing and the gatekeeping boards is some of the worst of the lot.
Most of all, because it actively interferes with the transfer of knowledge and serves to penalize those without means, and those in poorer countries.
Considering the vastly exploitative system and the outrageous prices and robber baron profits these industries have made...
I consider it my duty to accelerate its creative destruction and return the ability to recommunicate ideas back to its inherent social contract.
Using your bits and bytes to store data is a revolutionary act that this vampire industry of ghouls seeks to prevent.
Down with the literary Robber Barons!
I don't think it's the same thing for books by individual authors though. Good writing is a lot of work, whether it's bytes in a text file or ink on wood pulp. Giving authors a way to make a living at it so they don't have to wait tables the other 8 hours a day would be nice.
If publishers, agents, etc. are providing a real service to the author - marketing, networking, getting good work seen - then I don't see anything wrong with them getting a cut, if the author chooses to hire them. Or if the author wants to take that on, build a social media following, beat the pavement doing book tours etc. that's fine too.
But what's the point in doing this for legal content? You can just get the Illiad from the Gutenberg Project.
Some of those book even are in their own github project, but it would be nice to have them somewhere, and maybe in a way that could resist censorship o blocking.
Or other things than books, like projects or research released in public that some commercial platform centralizes and put a price tag to access them.
It has become substantially cheaper to do this sort of crowdfunded guerrilla knowledge sharing.
I’m not making any predictions about how long Z library stays up, but the illegal seeding of movies and tv shows has remained very strong until today.
Last time I tried it, the ipfs service used its own storage scheme. Meaning it's not like pointing Apache at a directory. You take your stuff, and upload it into ipfsd first, and it puts that data into its storage system.
So to do this from scratch (not mirroring somebody else's content) needs a minimum of >62TB -- 31TB of content, which ipfsd will then package into 31TB more + overhead in its storage area.
And of course if you're doing this, you're expecting other people to mirror this stuff, so count on hundreds of terabytes of traffic.
So this is easily ~$3K in hard disks alone, plus the NAS/server hardware, plus traffic, plus the willingness to risk the FBI coming and grabbing all of it.
It maybe helpful to leave the files in the file system as is, and store the metadata in a sqlite database.
Then integrate it with a p2p network layer for crowd seeding and even content discovery.
> So to do this from scratch (not mirroring somebody else's content) needs a minimum of >62TB -- 31TB of content, which ipfsd will then package into 31TB more + overhead in its storage area.
IPFS has `nocopy` option for quite some time now, which avoids copying.
> And of course if you're doing this, you're expecting other people to mirror this stuff, so count on hundreds of terabytes of traffic.
Of course you are expecting other people to mirror this stuff, and naturally that will generate some traffic. How is this not a problem with web mirrors or torrents?
Oh, didn't find about that one. Thanks!
> Of course you are expecting other people to mirror this stuff, and naturally that will generate some traffic. How is this not a problem with web mirrors or torrents?
I mean, if I put a book archive on the web, I'd expect the vast majority of people to just grab whichever book they were interested in. Mirroring is a possibility, but a non-trivial thing to accomplish, and can be discouraged.
Meanwhile, on IPFS I'd expect a much higher likelihood of somebody trying to replicate the whole archive, so one would do well to keep that in mind and to be prepared for it.
Re #2: Perhaps that's true, but on the other hand, the load will be distributed across all seeders with IPFS whereas your web server will be the only one shouldering it.
This is a very good argument for primarily using ipfs to host high quality content that people will want to preserve long-term…
there is a free service that does this for you till like 5 gigabits, it pins with filecoin
I have entire sites hosted for free this way, whereas other hosts like vercel and netlify will charge you for traffic. you can just put your big assets on ipfs+filecoin pins and have unlimited traffic. the ipfs CDNs help with performance.
I imagine the client could request for data to become available, then server (person managing a part of shadow library) could make somehow available the files after a certain period, and the client would automatically download the available files.
What’s stopping the FBI (or others) from taking down such linked websites? Isn’t just a matter of time until all of them are taken down and the owners sued?
How about you? How hard was your life destroyed by book piracy (or any other kind)?
Yes
Why can't I, a speaker in language X, buy books in that language outside of the country where this language is spoken on the same platform where it's sold.
See e.g. availability of, say, Japanese or Russian books on Apple Books in Europe, US, Russia and Japan.
But at the end of the day, I wrote my book to communicate knowledge I had and thought was valuable, and if people who couldn't afford the $30 got some value out of it then it was worth it. Now obviously it would be different for someone trying to make a living out of their writing.
Be sure that some people in fact paid you the $30 only because they tried it first and said "this is worth it" (it is how it sometimes actually works).
If it is to calculate missed revenue: are you considering that the other side of what I have written - a fact - is that a number of people who downloaded an available copy of your book would not have bought it anyway?
Those relevant to the missed revenue are those who would have paid but did not: the number in that group is something that you can hardly quantify, and on the other hand you have gained buyers that would not have bought your book if they had not tried it.
There's also no difference between the game downloader and them playing a copy at their friend's house and deciding not to buy it. There's also no difference if they were to buy the game second hand.
In a study commissioned by Google: https://www.ivir.nl/publicaties/download/Global-Online-Pirac...
Certainly not: a thief removes property, leaving the victim without the re-appropriated.
I think it’s a downright silly concept to pitch spending public money on cops to hunt down archivists rather than spending money on archivists. I think only the most embarrassingly naive folks would advocate for the world to work that way.
Imagine being so terrified of an imaginary John Galt scenario that you actually put people in jail to avoid the very thought of it.
<https://news.ycombinator.com/item?id=33647625>
Criminalisation of digital distribution was only legislated in 2008 (again in the U.S.):
While we are enumerating things we haven't seen...
In the United States, copyright exists "to promote the progress of the useful arts and sciences", yet copyright is defined in terms of "life of the creator plus some years".
I have never seen an explanation of how copyright is expected to incentivize a dead creator to perform further creative acts.
I don't want anything to do with anything illegal.
Unless you mean that you do not want to have anything to do with hammers and screwdrivers.
Today, thousands of elderly people will be scammed by idiots that spend their entire day preying on older people. Elderly people that will lose everything and end up homeless.
Fuck the current implementation of gift cards. They cause more problems than they fix. We can't have nice things because there are people out there who are fucking disgusting. Therefore we need accountability built into each technology.
Why? because there's a lot of people that fucking suck that's why. Build a simple application that allows people to share plain text and figure out what happens next. Harassment, death threats, organized crime, and all sorts of horrible things.
Technocrats and tech enthusiasts are often too idealistic to think about the real world consequences of technology.
The person that was having fun implementing convolutional neural networks and decided to publish their research for the benefit of the world would never have imagined that years later the chinese would take that research and build a mass surveillance system that tracked the entire population in real time enslaving billions of people.
The people that built Tor would never have imagined the caliber of fucked up shit that now takes place there. And the same with every technology.
It depends on what you mean with that «accountability built into the technology».
Related case, as you are also speaking of cards: this friend of mine wanted to buy a foreign book (less available in bookstores, where you pay cash). So he purchased an anonymous pre-paid credit card. He found out that during the mess of the past three years, somebody has made it very difficult to keep those cards anonymous: they now request a telephone number to activate them, and in regions where, "coincidentally", telephone numbers are to be associated to legal identities.
Now: as you should know, you can bet your life that if our transactions were recorded, we would strictly avoid commerce.
(Just like, pretty similarly, we would strictly avoid a non anonymous Internet.)
So: your call for accountability should work within the boundaries of what is acceptable.
> too idealistic
Don't worry, we grow up, and some of us restrain production - exactly in the awareness of what may happen.
Imagine working all your life so that then everything gets taken away by some jerk.
So if you run a node and want to avoid illegal content, the solution is just not to use the system to pin such content.