Pirate Library Mirror: Preserving 7TB of books (that are not in Libgen)
pilimi.org
pilimi.org
"Hey, wait, literally everyone could have the entire library of Alexandria in their house for a couple hundred bucks per person. Like, all the knowledge ever. Maybe that should be considered the good default of things.
At least one in every town that everyone could use, for free, forever, without restriction to ANY of the knowledge anyone desires."
A web server can run just fine on e.g. Raspberry Pi Zero W ($10), exposing any such content to any smartphone etc able to connect to it via WiFi (Kiwix sells preconfigured SD cards for their content, even). So, assuming that most people already have a phone or a laptop or even something like a Kindle, the only non-negligible cost here is storage. And a 1 Tb SD card can be had for under $150 right now.
So if anything, I think OP is overly conservative, given that their estimate was "$200 per person". Unless that counts the devices used to consume the content, and not just storage and the server.
An open source project could make a "search-engine-language-model" and in turn make this library much more accessible.
A recent article related to this: http://mitchgordon.me/ml/2022/07/01/retro-is-blazing.html
From the perspective of how much you individually could read in a lifetime: that's about 4,000 weeks (to roughly age 80-ish, and presuming you're not already reading at birth). Multiply that by the number of books you plan on reading a week. You'll probably read fewer than 4,000 books over your entire life. Many people read few if any books after graduating secondary school, and even a highly-motivated reader might be challenged to crack, say, 40,000 (ten per week for life).
The US Library of Congress has the world's largest book collection at about 40 million distinct titles.[1] There are another 130 million total catalogued items, including photographs, films, audio recordings, maps, pamphlets, and other items. Not all are textual. I'll stick to books.
At 4,000 per lifetime, you'd need 10,000 people simply to read all of the LoC's collection at a reasonable pace.
(By comparison, at the turn of the 20th century, the LoC's annual report to Congress noted that its cataloguing department could handle about 3,000 books per cataloguer per year, or a pace of about 60/week or 15/day. Over a 40-year career, a cataloguer might handle, if not read completely, 120,000 books.)
As a rough rule of thumb, a digitised ebook in PDF format runs about 5 MB.
Every single book in the Library of Congress in digital format would occupy about 200 TB of storage.
Current disk prices are running about $5 -- $20 / TB.
200 TB in raw disk would set you back $1,000 -- $4,000. Figure 2-4x multiplier for a disk storage system all told, and it's still roughly $4k -- $16k to have as local storage every last book in the US Library of Congress. That's well within scope of a moderately wealthy US household budget, and would be reasonable to consider for a small-town city library.
That's the technical storage cost, obviously not the rights or aquisition costs for the materials.
The Library of Congress's budget is about $800 million/yr.
Those 4,000 books you might read in a lifetime? They'd fit on about 20 GB worth of disk. If you're a 10x reader, 200 GB, and a truly dedicated 100x reader, about 2 TB. Those are well within the range of present desktop / laptop disk allocations, and represent a few tens of dollars of storage expense.
The Internet Archive computes $2/GB for storage in perpetuity. That's roughly 400 books worth of storage.
But a small household NAS and a very modest server platform (most routers can effectively operate as media servers) could rival a mid-sized city library for about $100 or less in actual hardware outlay. Those prices are falling by half about every 18 to 36 months, as they have been for decades.
Books are not all published content, and all published content is not all human knowledge. But standard published books are a good proxy for total cultural knowledge, and in all likelihood, then some.
________________________________
Notes:
1. See: https://www.loc.gov/about/general-information/ That includes about 25 million books in the main collection, and 15 million in "nonclassified print collections, including books in large type and raised characters, incunabula (books printed before 1501), monographs and serials, music, bound newspapers, pamphlets, technical reports and other printed material". I suspect you could roughly halve my estimates above based on 25 million vs. 40 million volumes, as many of the nonclassified works may duplicate the main collection.
Not to mention the cost of development - for that they needed thousands of GPUs or TPUs and a large team of top AI talent which is very very expensive to hire.
Regarding the open source copyrights - wasn't that code released to be "open"? This is actually a radical way to open the code, make it more useful for everyone. Open source devs can also use it to create projects.
I think $10 per month for this whole process is justified. It's not like you could run it at home.
I guess YC reaction to codex would be substantially different if this was a free and open tool anyone could use....
The idea that we do not do a world library of free digital copies of every book ever written really highlights the problem with the thinking your comment has demonstrated. The idea that individual pursuits must make a profit to be justified leads to us doing terrible things like: not making a free world library of every book.
Though in this particular case this also demonstrates a major problem with intellectual property concepts. Actually hosting the library isn't very expensive. But we have made doing so illegal. Of course, authors deserve to live a decent life just like everyone else. We currently do that by restricting all access to duplications of information they have produced so that they can charge a fee for access, and that fee provides for their survival.
But we suffer an incalculable loss by making all this information restricted. In my view we would be MUCH better off as a society with respect to creativity, innovation, and other popular metrics for progress, if we actually made sure as a society that every person's survival was provided for with no need for them to pay for it. Then authors wouldn't need to get paid, engineers could do what they love to do and post all their work as open source, and we could have a free library for everyone. This extreme openness would in my mind lead to more rapid innovation, and markets would still function as first movers would maintain an advantage for new product releases, though they would have to keep moving as anything they've done that is worthwhile would be copied. But since no one's livelihood would be at stake, this is not a real issue.
This can all be done in a voluntary, libertarian society as long as we have community ownership of the means of production, and promote these ideals of community support in this society. And I think we would be way better off. Doing this would allow us to offer every book ever recorded for free to every person on Earth. A big change, but one with obviously a very big benefit to humanity.
One note though: people who want to own a lot for themselves really mess this up. So people would need to dissuade those people from acting that way. My preferred method of doing so is by starving them of workers and customers, though when it comes to control of land matters get more serious.
I do not do that anymore. The world is clearly full of cases for which absurdities are tenable - not by an odd minority, but by "(pseudo)random people". If you declare them, the presumption of irony is gone.
Which is not to say that I approve, merely to state the root explaination for the insane state of affairs. You could say dead serious in the sense that I wasn't joking, that really is the explaination.
I don't say I like it or that it's reasonable or fair, it wasn't approving.
It's not sarcastic because it's literally the actual explaination.
Libertarianism includes the right to own property.
Also common ownership of the means of production is an old idea, and has been tried many times. It always results in poverty.
I’m not saying people shouldn’t have the right to own property. But that a good way of organizing society is collective ownership of the means of production. If you are part owner in something with shares and a contract, that obviously still relies on property rights. That is how the stock market works after all.
EDIT: I am basically proposing a change in norms, rather than a change in rights. Currently the norm is individual private owners or ownership by a board of directors. I am proposing ownership by communities as a collective. Same rights involved, but a different norm.
I suggest you contact all like-minded authors and suggest they join your book collective and donate all their intellectual property rights.
And we can certainly discuss the merits of our government granting ever increasing copyright terms. Copyright is not a natural concept, but an artificial one created by government agents. Surely this forum is a good place to question those choices and discuss alternatives.
Have you heard from any?
> copyright terms
I have a history here of advocating short copyright terms. I speak as someone who made a living selling copyrighted materials. The stuff I write today is all released under the Boost License, which is as close to public domain as it gets (because some countries have no legal concept of public domain).
I also have a history here of dispensing with patent laws, even though I hold some patents.
Thank you for the clarification. When I try to discuss these ideas I feel like you respond, without asking what I have been doing to pursue this, by saying "go and start one then" "it's not illegal" "go find other people to talk to about this then", which feels very much like you do not think that is exactly what I am doing when discussing it here. It feels like you think I need to be somewhere else to be doing that. And it feels rude that you tell me to go and start one without understanding that I am laying the groundwork in my own life to be able to do something unprofitable indefinitely (my collective will be a non-profit collective). I cannot just quit my job and do that tomorrow, but I have been working very hard for years to arrange my life appropriately to follow that pursuit, and I have made great progress. Already I am getting several hundred dollars in donations to my non profit. Not enough to live off of, but I continue to produce youtube videos promoting my non profit engineering. Your comment of "go and start one then" feels very dismissive.
> Have you heard from any?
From like minded people? Yes! My generation is extremely interested in libertarian leftist ideas. In providing for everyone regardless of the market value of their skills. I run two small youtube channels where I promote these ideas, though there are many very large channels with millions of subscribers who talk about them with more skill than I can, as most of my time is dedicated to my non profit engineering organization.
I am glad we agree on copyright and patents. I too have a few patents but would prefer we dispense with patent laws. And I release everything of mine under permissive open source licenses.
> you want to argue without really thinking about what I am saying
I've thought about it for many decades. I'm not trying to suppress your discussion, ideas, or any efforts you may make towards creating a (voluntary) collective.
Feel free.
BTW, I read somewhere that over 10,000 collectives had been created in the United States. Where are they now? Oh, they all failed.
A famous one was the Summer of Love in San Francisco in the 1960s. It only lasted for a summer. Seattle's Summer of Love a couple years ago lasted a few days before imploding. (Google CHAZ - Capitol Hill Autonomous Zone)
It really feels like you are not reading my comments. I just said in the comment you are responding to that I am working on this, but it is a few years away. That is not to say I am not doing anything now, but I have to lay the ground work. For me to say I am actively working on this and to read your response as "go ahead and do it then" really feels like you do not want to engage with what I am saying.
> BTW, I read somewhere that over 10,000 collectives had been created in the United States. Where are they now? Oh, they all failed.
They all failed? Walter, I can easily find a long list of active co-ops just in the Bay Area alone. It seems to me you are speaking with authority on a topic you don't know very much about. https://www.cooperationrichmond.org/coop-movement/worker-coo...
That said, failures do happen. Lots of privately owned businesses fail too. Would you suggest then that privately owned business as a concept is not viable?
I think what you are missing is that people in today's economy are suffering, and people are desperate for something to change. I suspect you are doing okay, and not desperate for change. Good for you. But the number of people in your position is literally dwindling from a statistical perspective in the USA. For everyone else, they cannot just sit comfortably and dismiss suggestions on how to change things. They need change, as for many it is literally a life and death situation.
A long list? A handful out of the zillions of companies in the bay area. It's statistical significance is zero. Some are charities, which are obviously not self-sustaining. I wouldn't be surprised if others were sustained by government checks. People start collectives in the US all the time, but the test is if they last. A typical collective seems to last about 2 years before imploding.
> Lots of privately owned businesses fail too. Would you suggest then that privately owned business as a concept is not viable?
Compare your list of a few communes with what's in the Yellow Pages for the area for the answer.
> I am working on this, but it is a few years away
While I wish you success in your endeavor, it's not really a collective unless you have other collective members signed up and collaborating with you. From much personal experience, I can attest that announcing one is working on something doesn't mean anything to anybody. It only matters if you've got something to deliver. I know that sounds harsh, and it is, but that's how things work.
> what you are missing is that people in today's economy are suffering, and people are desperate for something to change
I understand very well that they are suffering. They are suffering as the result of leftist government policies, not capitalism.
You are moving the goal posts. You said "they all failed" but this is false. I was pointing out your error.
> Compare your list of a few communes with what's in the Yellow Pages for the area for the answer.
I believe from the above two quotes you are saying that co-ops have failed as a concept on their own merits. I.E. we can observe their relative popularity compared to hierarchical corporations and conclude that co-ops as an idea are a failure.
But of course, as you seem to be a libertarian capitalist, you must also believe in the correctness of ideas which have so far failed to gain traction. In fact there are many philosophical ideas throughout history which were superior to the status quo but had not gained traction. For example take democracy versus monarchy. Would you have looked at Europe 1000 years ago and concluded that democracy had failed in the marketplace of ideas? Certainly not. In the same way, cooperatively owned businesses are part of a broader labor movement which was systematically attacked by those in power to dismantle its strength.
One notable example is the Taft–Hartley act of 1947, which placed significant limitations on the legal right to strike or boycott. I am not an expert on the labor movement in the USA but without going in to an entire debate on the subject, you can understand how the existence of unions and cooperatives (which I lump together as part of a common labor movement) may not have failed purely on its own merits, but may have failed due to attacks from organized special interests. So without proving or disproving that claim, you can see how a simple examination of the state of the labor movement today is insufficient to conclude the merits of those ideas as they relate to the average person. Just as you could not look at Europe in 1000AD and conclude that democracy was an unworkable ideal, and that monarchy was obviously superior.
> I understand very well that they are suffering. They are suffering as the result of leftist government policies, not capitalism.
You mean the leftist government which placed significant limitations on the power of labor organizations? The leftist government which funds the largest war machine in human history? The leftist government that remains the only major country in the world without some form of universal health care? Sorry Walter, but only in the wild fantasies of FOX News commentators looking to get re-hired and Republican politicians trying to scare their base to secure their election is the USA a leftist government. Though to be clear, I am a libertarian communist and I want the US government to stay out of my way. This is why I advocate for cooperative ownership of the means of production and not government control of industry.
I appreciate that you finally chose to engage with what I was saying instead of dropping the same tired arguments I have heard and disproved so many times before. But of course I still disagree with your position. My issue with capitalism is not free markets. My issue with capitalism is direct control of our economy by an elite few who have class interests that run counter to the interests of the average person. The mismatch in interests between those in control (who seek power and profit) and the other 99% of the world population is what creates poverty and strife. And only when the people have control over the machinery that provides for their survival can we truly be free.
Failing and disappearing is a feature. Failing and lingering on through violence and coercion is the real sin.
That said, I think privately- and collectively-owned enterprises can coexist (and compete!) just fine, and in a truly free market we'd see a healthy amount of both.
In the US, at least 10,000 collectives have been formed. You can form one any time you like. They're not illegal.
I think we would be better off if we do this with movies, music, games, and engineering designs like designs for medical equipment, cars, etc. most of the standard economic arguments for intellectual property restrictions are wrong and short sighted, and I have heard them all. I think this concept is well worth exploring both philosophically and in practice
The broader class of cultural products includes different types of expression as subclasses: the discriminator is in the broader class - cultural products.
Who benefits? Who loses? And to what extent?
And what would the net total cost of rectifying any shift of advantage to a copyright-minmalist regime be?
For that last, look to the net total annual revenues of commercial media.
And keep in mind that we're carrying out this discussion on a technological ecosystem founded in very large part on FSF Free Software / Open Source software.
That's about as relevant as saying we should keep in mind we are carrying out this conversation on hardware mostly made in China (and most free software is also written on hardware made in China) so perhaps that is where we should be looking for answers.
Your comment above suggested that stripping direct profit motive and incentives from copyright would produce a net harm. The Free Software model is based on the premise that stripping most exclusive rights under copyright, and specifically the ability to create new copies of, and new derivatives of works, is actually a net benefit to society, and often to the original author of a work who benefits by the collaborative and collective development and advancement of it.
One of the more fascinating characteristics of not only the Web but modern computing is that it's been the abandonment of proprietary interest in works (code, standards, protocols, operating systems) which have delivered the greatest value. AT&T did not voluntarily relinquish UNIX to the world, it was forced to do so under a 1950s US antitrust decree which granted the firm its monopoly in telecommunications but forbade it from engaging in the computer business.
https://www.softpanorama.org/History/Unix/unix_chonology.sht...
Arpanet, IETF, the GNU Project, TBL's CERN document-distribution project, Linux, multiple programming languages, and other elements are all based on the notion of non-proprietary ownership rights to the extent that free use and (generally) modification are permitted, often with little further restriction (MIT/BSD licenses), or with the requirement that further distribution also requires source availability in preferred form for making modifications (GNU GPL, AGPL, etc.).
That is: my comment was salient to the (implied) central fallacy of your previous statement. A foundation which your subsequent replies suggest is in fact your point.
no one worth agreeing with starts a sentence with >Maybe if you read the right books
It's not what most people would want when given a choice, but that does not mean it's very hard to imagine how and why it happens.
I did not mean to imply I approved either, just that, I think "this is insane" is kind of silly. It is, but, it's also inevitable and not new or hard to follow.
In our current world, the currency is money, (vs land or humans in the past) and things pretty much only get done if someone can make money from it, and someone else doesn't make money inhibiting it. That's it. That's the entire mystery.
The unspoken answers to my two questions were "no one" and "many many many". Everything that you can do for yourself, some company somewhere would rather you pay them for it instead. The product doesn't exist because there are zero companies out there with any interest in producing it, and unlike pure information (software), individuals can't just do it for free for the feel goods.
It would require some imaginary and highly different societal structure to alter that equation. The people who would manufacture the home libraries, where do they get the materials and facilities from, and how do they eat, if they aren't selling these either to you directly, or by way of your government and it's taxes? Some kind of co-op type organization that takes the place of a for profit company? And they get their electronics parts and their vegetables from other co-ops since they don't have money to buy them? It could all be done some other way than by counting dollars, but literally everyone in the world would also have to be doing things this other way. And all the people currently on top of the current system will not be the ones rewarded with success under any other kind of system, and so they wield their power to keep things exactly as they are.
I don't mean in an illuminati way like there are 11 people running everything, I mean countless people inhabiting all levels and niches all act to preserve what they think is all they have, in countless little ways. Even people who actually have essentially nothing and would live far better in any other system, because it just takes more imagination than most people have to even consider not having dollars, and more generosity of spirit than most people have to even consider not being able to boss other people through their need of dollars.
We would somehow have to collectively figure out how not to reward the very worst of us with all the leadership positions, both private and public. Ultimately it comes from that. The exact wrong kind of people to make rules are just about the only people who gun for those roles. There are all kinds of asshole things I wouldn't do if I were running things. And I have no interest in being anyone else's boss, nor in figuring out ways to parasite off everyone else or harness them to some will of my own or anything like that, and so I will never end up running things. You have to be a sociopath to do what it takes to ever get into such a position. It doesn't happen by any nice fair reward-the-good kind of way, and we have almost no mechanisms to detect and defer such people from suceeding in their shark-like process. It just works and they go right to the top.
I can guarantee you that this type of behavior means I'm writing fewer books. It's very short-sighted.
Historically most art was financed through patronage/sponsors, wealthy people or institutions that sponsored people for the social status, or commissioned works from them. Both of those systems have been democratized by the internet, with people sponsoring artists they like with systems like Patreon, or just straight paying them money to create something. The same is happening with Twitch subscriptions (essentially donations to the streamer). There's no reason it can't work on a larger scale for most literature. Sure, priorities would shift, but it might also enable authors more creative freedom than publishers are comfortable with.
*In general, so quite possibly also yours.
Everybody deserves to be paid for their work and I would argue that this is more about preserving knowledge than it is about not having the writers get what they deserve.
Yes obviously pure capitalism is as shit as pure anything else, and no system made by humans, and made of humans, actually works very well to create sanity and fairness. So even not-pure capitalism creates all kinds of undesirable results, like available tech doesn't get used in ways that would be wonderful for everyone.
What solution to human nature have you imagined in your split second?
The point you’re trying to make (I think) is that information isn’t particularly useful on its own. True. It has to be relevant. With good, useful information you still need community support (whether that’s your household or your country), funding, time, etc. But that’s going to be the case whether you’re using a book or learning from a forum, or a video. Internet videos are worthless if you have no internet. Tribal knowledge is useless if no one can get the information from the previous generation. The medium is not the problem, and suggesting that relevancy is a fault of the medium is a misrepresentation.
Germany now is a leading copyright-maximalist power, unfortunately.
________________________________
Notes:
1. I say nation and not country as the German language and cultural nationality pre-existed the German state in 1871.
I still can't understand why can't every public library host media that should be free and easy to get.
It's a nice sentiment but like, people can already go to gutenberg.org and download pretty much most important works of literature in existence and most books have like 5k downloads so there's that.
(Also, keep in mind that content is now entering the public domain every year, and projects like PG are nowhere close to keeping up with that flow of newly-unrestricted stuff. So this dynamic is becoming more extreme over time, not less.)
> The overwhelming majority of our knowledge is in the public domain
https://cand.pglaf.org/germany/index.html
Gutenberg was blocked in Italy in early 2020. I don't know if this block still persists.
https://www.balcanicaucaso.org/eng/Areas/Italy/Project-Guten...
afaict, it resolves and loads fine here. Now on mobile tethering via the Iliad (.it) local operator. When I get home I'll also check it on broadband.
I wonder how your provider overrides it. Is it too much of a wet dream to think that some operators (behind the DNS) responded to the requests of the vacuously appointed low scum* with disdainful neglect? In terms of "If you really had any credential for existence but some forms of physical expression, we would say that you must be joking - before putting the proof of your vilty in the trash archival it deserves".
*May I remind that they appended the apex of Civilization, Project Gutenberg, to a list of a dozen pirate sites in their "operation". No, really, I cannot find words to label those lowest abysses.
I don't use the (ex) national provider (Telecom Italia / Tim) however. Crappy service aside, they're well known for jumping every time the government tells them to, and happily filter a lot of stuff.
- Gutenberg - Zlib - libgen
via o2 (not promoting them in any way)
By copyright expiry, public domain in the US begins in 1927. Later works may be in the public domain, but all works published prior to 1927 are in the public domain in the United States.
(This may not be the case in other countries.)
There were not many published books prior to the invention of the printing press, and many of those didn't survive. The total number of books (not individual titles but actualy bound volumes entirely) in Western Europe as of 1400 may have been as few as 50,000.
By 1800, about 1 million titles had been printed.
Over the course of the 18th century, presses became vastly faster, as they evolved from hand-operated wooden screw-press to iron-frames to steam and electric-powered rotary and ultimately web presses. Paper became much cheaper (and less durable --- a factor commented on at length in the Librarian of Congress's annual reports to Congress in the late 19th century). Literacy exploded from ~25% to 95%+ over the 19th century (and probably accounted for numerous revolutions and political upheavals).
Through much of the 20th century, certainly by 1950, US publishers were issuing about 300,000 new titles per year, a rate which state remarkably constant through the early 21st century. By the aughts, "nontraditional" self-publishing (a/k/a "vanity press") was nearing or exceeding 1 million titles per year more than had been published through all time to 1800.
Reports that all recorded data was doubling every few years date to at least the 1960s. That would mean that in any two year period ... half of all recorded information was less than two years old.
The catch is that not all recorded data is published. So I'm not sure what the time-distribution of all publishing looks like. But I'm pretty confident it's skewed far more recently than 1927. And would thus tend to be copyrighted rather than uncopyrighted.
If you want to measure works by significance, you might make a different argument --- there are many great works of literature, philosophy, history, and religion which were first published before 1927. But ranking and tabulating these is more challenging than a simple enumeration.
There are projects out there that lean in the direction of offline viewing of lots of content, for example, having an offline backup of Wikipedia, such as:
- https://wiki.kiwix.org/wiki/Main_Page
- http://xowa.org/ (HTTPS seems not to work though)
I just wish that the process of actually accessing the data was a little bit more straightforward: https://dumps.wikimedia.org/backup-index.html (given how many different files there are to choose from, you'll probably want to read a tutorial or two)That said, while text is perfectly doable, things do tend to get more difficult if you also want images or videos, because those do take up a lot of space.
We finally build the tools to allow the free sharing of information and almost immediately a bunch of rent seeking lawyers lobby congress to make doing that illegal.
I've always lived by the rule of "If you put it on the Internet, it's not yours anymore." Because that's de facto the truth. Only if you employ lawyers do you have even the slightest real power over something online, even if you're the original source. This isn't something 95% people can afford. Turning the Internet into just another place where the rich get preferential treatment is a terrible thing.
Go to your local library. There is actually a lot more there than books. There are videos and audio records too. There are news paper clippings too.
Worse, data is being generated very fast. Let's look[0]
> [In 2012], CenturyLink projects that 1.8 zettabytes of data will be created. By 2015, the projection is 7.9 zettabytes.
> MAST is currently home to an estimated *200 terabytes of data*, which… is nearly the same amount of information contained in *the U.S. Library of Congress.*
But that said, a few petabytes is going to only cost in the 10s of thousands of dollars and max in the low 100s. So this is well within the budget of even every modestly sized city and definitely every university. It would be even easier if this was done with torrenting.
I think the real question is^: why are we not, as a species, creating this system? It seems very reasonable that we could back up all data. (There is also a dark side to this too! So let's not forget about that) At least, why aren't we creating a full access world library for all scholarly data, which is probably something that could be housed for a few grand. Should we do this on our own given the relatively low cost? Books + sci-hub + arxiv + *xiv?
[0] https://blogs.loc.gov/thesignal/2012/04/a-library-of-congres...
^ maybe it is already being done and I don't know, please inform me if it is. (Other than the NSA)
So, why not just re-upload them to Libgen, then? I guess somebody will do that now anyway, but you could easily done it in the first place, without making your own mirror, which is not a mirror of Libgen. Just upload them to Libgen and make a mirror of Libgen.
> Q: Should the Z-Library collection be added to Library Genesis?
> A: Yes! However, it is tricky. Library Genesis splits out its collection between non-fiction and fiction. They also have relatively high quality standards. If you are interested in organizing all the books to meet their requirements, let us know.
I OCR'ed it, and I'm slowly fixing the errors and converting to epub.
Tedious, but interesting.
However, I also contacted the author, and sent them a copy! At first, they were furious - I'd even translated ('pirated') the copyright notice...
But things settled down, and now I'm working with the author on translations to various other languages. It's given the book a whole new audience!
Put it in a case with all parts, $800USD all up?
But it is also use to protect unique creator revenue and encourage to create more.
If you ask where the fine line should be I have no immediate answer, but abolishing intellectual property rights just like enforcing them at all costs doesn't seem to be the optimal course of action to me.
And this arguably should be extended to tangible assets as well - I like Singapore model where housing property is sold for specific timeframe. It simplifies a lot of redevelopment.
We tend to think of ownership as in absolute owning of an asset for indefinite time. For a lot of things in Singapore you can ownership (i.e. own it) in that sense, but after sone time you have either to return it or stop using (rendering it useless). This applies to assets like homes, cars, etc.
If it’s not feasible to enforce those policies, government just imposes hefty tax on assets, ensuring you extract (or contribute) sufficient added value from asset.
This thinking is an artifact of an economic system so dependent on scarcity for its motivation that it is now generating most of the scarcity in the world.
We now have the technology for creative implementations of "From each according to its ability, to each according to its needs". Just keep track of how much each thing is used, and reward creators from a corporate-tax-funded pool. Every for-profit entity contributes proportionally to its profit, and can use any idea for free.
Should I be disallowed to commercialise it?
I partly get where you stand but if I was in a society that you seem to endorse my first question would be, other than for the love of doing it, why sink so much effort into a thing only to get nothing back. It almost is the opposite of a meritocracy.
Note that if you don't defend it in court, our justice system thinks it is less valueable to protect. Which in itself is kind of ridiculous.
Also, right to commercialization has nothing to do with intellectual property.
There's always someone who just has to spread the despair and helplessness. Every bloody time. Tell me, in what way does this add to the discussion? I wish there was a ban on these kind of comments.
> right to commercialization has nothing to do with intellectual property.
I don't understand. If I don't own it I can't market it, right?
It you only want answers you like go talk to a mirror. The way this adds to the discussion should be pretty obvious, but let me spill it out: If one of the main arguments for IP is wrong it's costs/benefits have to be reevaluated.
IOW he didn't say what you claim. Plus he gave no way forward to achieve his goals, and am I not in any position right now to Bring Down The Man, much as The Man may need it, so it was just a hopeless valueless post.
As such, I would say what they said is absolutely what I claimed. I just explained it in plainer terms and without requiring you to (re)read the rest of the thread.
Edit: The fact that you summarised "this system must be destroyed" as "hopeless" says something about your own fundamental hopelessness and despair. Which you kind of ironically attributed to someone else.
I don't agree, just the excessive enforcement of it may well be.
> The fact that you summarised "this system must be destroyed" as "hopeless" says something
No. What I said was:
> hopeless valueless post.
The post was valueless because it gives no direction, no means. And I detest such posts because they offer nothing useful. They are unconstructive. Hence are valueless.
ONLY private persons should own patents, and it should be illegal even for employers or institutions (academic, research) to own patents of people employed to do research. At most companies and istitutions should be allowed to add a clause of "perpetual-free-usage of any patents of employees resulting from direct work" - but an employee or group-of-employees holdig a patent should still be able to license it to other companies too. If businesses are hurt, that's GOOD, most should not exist as coagulated entities.
We're not gonna have proper freedom preserving capitalism ultil we properly decentralize: we all work like swarms of 1-person-companies / solopreneurs contracting between eachother. (No, not the gig-economy, in that distopia we're all still slaves that can't band together to fight the masters.) Legislation will automatically have to be refacored to make this work. With some exceptions, only human individuals should hold most property, not companies and not institutions. Groups/collectives only when the group members directly worked together and know eachother.
And Intellectual Property would just "click in" in in such context. IP sounds hellish and disfunctional because our own practically techno-communist society (yeah, even USA is practically "communist" nowadays in a way - newsflash: "the reds" have won! even the f symbolism is there, "the red pill" is the good one now... all's backwards) is messed up. It makes perfect sense in a hyper-decentralized hyper-individualistic REALLY democratic and REALLY capitalist society.
Of course I like to be credited for my work — but some fan adding my work to a pirate page would not be a concern but rather a bit flattering. What would anger me would be someone claiming credit for themselves, some rich company taking the material without paying me, things of that sort.
Nothing wrong with expecting to get paid for your work. If as a consumer you don’t want to pay, stick to open source and freely licensed media.
There are quite some hoops small self publishers have to jump through to get their music sold in a way it can actually interfere with the big corps in that space. These hoops are all there to make the market entrance harder, they are not there to protect artists.
No one said anything about disallowing you from doing anything.
>my first question would be, other than for the love of doing it, why sink so much effort into a thing only to get nothing back.
Is it wise to sink your time into something you don't really like doing?
or it might. Your own answer acknowledges that with the 'might'. So yes, you measure the odds then throw the dice.
I believe you are mis-phrasing the question. What you're actually asking is:
> Should the state criminalize and punish people who make copies of my work, to facilitate my commercial activity with it?
And our answer is "No".
You can go ahead and engage in whatever commercial activity you like, based on open access to your work.
Overheard from an IP lawyer I stood near to once - something about f/oss software being incorporated into commercial products being a big issue (for the free stuff, not the company doing the 'stealing' of it). Your view?
the licences are being ignored. It's what she said.
You:
> This is the exact reason why copyleft licenses are important: you can reuse (A)GPL content, but if you do so the result must be given back to the community
Me: the frigging licences are being ignored. Giving back to the community is not happening. Code is being stolen. Are you trying to ignore what's being said?
Without copyright this situation is a lot more equalized as now you can have people make modified versions of macOS and redistribute those legally - yes, not having source access makes that more difficult, but not impossible and even if you did need the source, it only has to leak once.
Welcome to free software. :)
If you had to "pay back significantly" to use it, it would not be "free software".
It's regrettable, but it's the price of freedom. Whether that's worth it is subjective. Stallman's answer to this was the GPL. (I bet Apple wouldn't've touched BSD if it was GPL'd.) Newer, hybrid license have also emerged (like the MPL) that attempt to strike a better balance between freedom and back-contributions.
Can I freely use your toilet then? your electricity? You pay for the plumber, why don't you pay for entertainment?
Scientific knowledge must be open, but most copyrighted work is entertainment.
No because once you used them I don't have them anymore. There's a reason IP has different rules than physical property.
> You pay for the plumber, why don't you pay for entertainment?
Ok, that's a more appropriate analogy, but then again, should my plumber get a recurring fee for the work they already did, when I use the faucet to give drinks to my friends?
I mean entering your house, doing my business in your toilet and leave. I won't take your toilet with me, I'll come back when I need it again.
> when I use the faucet to give drinks to my friends?
But your friends all have their own house with their own plumbing they paid him for, so he can continue making a living from his craft. Writing a book takes months or years, not 2 hours like repairing a toilet, so of course the author needs to ask for money from everyone who wants to access it.
We're moving the goalpost here. And we're still talking about physical property (or possession) vs. intellectual property. I suggest we stop with that line of reasoning/metaphor.
> But your friends all have their own house with their own plumbing they paid him for, so he can continue making a living from his craft.
Again, the metaphor does not hold. Such a situation only means that the plumber is the only plumber in town. If we have multiple plumbers (so we can stick to the metaphor) my plumber can't forbid me to use my plumbing for certain uses (like watering my plants or offering water to my friends for free or for a fee).
> Writing a book takes months or years, not 2 hours like repairing a toilet, so of course the author needs to ask for money from everyone who wants to access it.
So it's just a quantitative difference? I can pay 1 cent per 1000 toilet flushes then. Seems fair.
My point is, I think these kind of metaphors don't work here precisely because intellectual work is its own thing.
If your use of my toilet doesn't affect me in anyway then I don't see why I should have a problem with that. If we are talking about you stinking up the place, using all my toilet paper and blocking the john whenever I need to go then we are talking about something very different from "IP".
You can freely use the design of my toilet, certainly.
> You pay for the plumber, why don't you pay for entertainment?
I pay for physical objects that are made for me, like books; and I pay when people come play their music (even if payment is not mandatory).
But TBH - I don't think that's the appropriate moral basis for things. For example, we don't pay for the huge amount of work our parents do for us; nor for the not-for-profit activities we often rely on etc. I would much rather support a non-exchange-based social arrangement.
EDIT: ach, didn't read your post propely - first line says " I'm also write technical writing (including academic publications)" so you have a strong position to hold your view - sorry
(deleted)
So did JK Rowling. So what's the difference?
I'm all for finding ways to reward people for creating things but that should not include restricting what others can create.
Plus fixing the plumbing usually only needs to be done infrequently, whereas a book is read in a comparatively short time, so from that point of view a consumer would also be willing (and able) to spend more per plumbing fix than per book.
So consequently you need some sort of arrangements that allow for splitting the necessary payment to the author up across multiple people and/or over time.
Additionally, artists often speculatively create works without knowing for sure whether the public will take any interest in their work, or not. Copyright certainly has its faults, but it does cater for precisely that scenario by ensuring that you can insist on getting paid afterwards if people enjoy and want access to your work, and you don't need to acquire all the necessary funding up front. If you can't come up with enough money, you can even "just" invest your spare time instead and still get paid back if the work turns out be successful.
Plumbers on the other hand I assume rarely have the desire to speculatively fix up other people's plumbing and then hope to get paid afterwards if they did a good job.
Rewarding artists after they've already produced the artistic work if they're successful also makes sense in that the quality of artistic output can vary, and so there's a bigger risk of disappointment if you need to pay far in advance, before the work has possibly even been produced.
And because the quality of an artistic work is also very much a subjective matter, it'd also be much more difficult getting your money back in that case, whereas plumbing can mostly be judged according to much more objective standards, so getting your money back – through the legal system if required - is again a more tenable affair.
Of course the existence of Kickstarter and the like or even just plain old pre-orders show that to some extent people are willing to take that risk of paying in advance, but whether that would be enough if it was the only reasonable source of funding for artistic works? It'd also mean that if you can't convince people to pay you in advance (and good luck with that if you're some unknown newcomer), then good luck getting any more money afterwards, even if the book/… then turns out to be wildly popular afterwards.
ofc not
but since noone can ever prove that his was the first incarnation of an idea, nobody can be criminalized for also doing things in a certain way.
the concept of protecting invention for some time to facilitate reward is not without merit, but the implementation of IP law and practise has gone so far astray that it's overdue to rethink the whole thing.
With grammar like that, it's probably just as well.
Writing is hard work. Writing books and getting them technically correct is expensive. This is very short-sighted.
Otherwise, I am having a really hard time understanding how can you suggest that I don't own the book I spend a *decade* to write. It is just as mine as the car you drive is yours.
To tighten regulations around intellectual property to make sure that it is not abused - sure.
To ban? Obviously never.
It could also result in far more draconian DRM, as that would be the only way left to protect your work.
Now drastically lowering the time of copyright might be well worth it, something in the realm of 20 years should be enough. As copyright needs to get back to a point where things you consumed in your lifetime, make it into the public domain in your lifetime.
There is no way to protect video, audio or text from being copied. DRM just prevents low effort consumer copying.
Or more practically, just look at cinemas. They already film the audience to prevent filming and with great success. While you still get illegal copies of a movie easily, it's only extremely low quality smartphone rubbish. The high quality piracy videos only shows up months later once the films hit streaming services or Bluray.
And all of that is just current tech, lets assume VR will become a success in the future. Now you have a device on your head that tracks every little one of your moves, including things like heart rate and eye-tracking. Furthermore, what streams to you isn't an easily ripable 2D copy of the movie, but the 3D view of sitting in a cinema. Good luck trying to rip that. And of course tamper proof hardware is a thing as well, so any attempt at opening it up will automatically self destruct it and phone home that you tampered with it.
Other than that, all DRM does it makes it harder, not impossible, to copy.
So, "you are being watched as you watch". Not by an uncaring attendant, but by some actively processing automation.
I think you've identified pretty much the only way (that I can think of, anyway). If you surveil all consumers and instantly arrest them the moment they make a copy, mission accomplished.
Short of that, though, as long as the data is being presented out in the open (light waves, sound waves, text), it's going to be possible to "rip" it.
You may see lower revenues, but how much cost is currently poured into resolving licensing / investing in DRM / etc, all for works to be pirated anyway?
So, when you have good content with reasonable prices, people also come and buy.
Also, there are some eBooks in Kobo store devoid of any DRM. So, publishers are not forced to use DRM on Kobo, as well.
God no. Right now, we'd be getting remakes from every piece of pop-culture that was semi-popular in the 80s-2002 time frame. Not just movies, but TV series, books, theatre, musicals, ...
Sure, copyright should be shortened, but I don't begrudge (eg.) a one hit winner making money off their hit decades layer. Life of artist is reasonable, I think. They take a gamble on a profession with risky pay-out; if it works out at least once for them, let them reap the benefits.
That's exactly what we are getting right now though. Looking at the top ten of the box office right now, only three are not part of an existing franchise. Three of them are reboots of 80s movies, four if you count comic books. Large IP holders recognize under the current system, it is much more profitable to exploit their existing IP than to come up with new concepts. If the copyright terms were significantly shorter, the pressure to be original would be far higher.
We are getting these things anyway, except that the originals are far less accessible than they should be. Entertainment trends are cyclical.
Put differently: we'd be inundated with much worse schlock than we get now.
It's not the job of copyright to allow people to get lazy or companies profiting forever from the rights they bought. The goal should be to encourage original works and current copyright isn't very good at doing so.
Also it's not like the author would go completely penniless here. Just because everybody can make a StarWars doesn't mean there won't still be a George Lucas approved canon-StarWars. Slapping the authors name on your product to declare it the "Read Thing™" might still be worth a bit and might frankly be better than today's sequels that happen completely without any of the original creators being involved.
Also, the information about nuclear, chemical and bio weapons should be accessible to everyone. Preferably as DIY recipes, that you can follow at home.
If a regular citizen can get his hands on tools/materials to make things like that, you are screwed sooner or later.
the system around intellectual property has some issues but some form of protection / ownership needs to be there.
if you had your wish and the concept of IP was treated as shunned and taboo you would quickly live in a world with vastly diminished amount and quality of art, science and technology.
But we live in an infant society, with many grown ups acting like spoilt childrens, saying “I want to get that fancy FAANG job, I want to be wealthy, and I expect to do it copy-pasting others people knowledge and infringing IP, but if someone else begs to differ I start whining”
1. there is some legitimate issue
2. people protest (peacefully)
3. a provocateur does something over the top (violence, absurd statements like "defund the police")
4. legitimate protesters are discredited because of 3.Anyway, "intellectual property" has proven to be a driver of quality, as the earned money gives liberty and time for the creators. I don't see how this is a bad thing. Sure, there are warts in the system and we should get rid of them, but not by removing the whole good side.
In that sense, these piracy sites are acting like global public libraries open to everybody with an internet connection.
At the same time, I feel the authors, researchers, editors, and other support staff that gift the world with knowledge should be rewarded for their effort.
It'd be great if there's an honor system that enables readers around the world to pay them some amount to show gratitude.
The current system is of two extremes -- either first pay the price set by the publisher to even browse a book (and that price is ridiculously high in underdeveloped countries), or get the full book without paying anything.
There should be a spectrum of rental and gratitude amounts in between. The publishers themselves can together set up such an online library to make it all legal. Not only will they help humanity, but they'll also get some of the revenue they're currently missing out on. A balance seems to have been struck in the music business with most of it being legal and accessible nowadays. They should do it for books too.
I want a piece of software to which I can add a collection of files, say multiple TB. The software will then behave a bit like a BitTorrent tracker, and know which peer has which files. A peer joining this swarm will be able to say "I want to donate X GB of space", and the tracker would tell it "OK, then download and seed these files, which are the least seeded".
The peer would download the files from the rest of the swarm and make them available to it. Then, a request layer on top of the swarm could be used to request a file from the peer which had it. Adding/removing files to this collection would also need to be a feature.
Does anyone know if anything like this exists? If not, how easy would it be to make something like it out of BitTorrent? I might give it a go.
It's not very convenient for archival, whereas the system I'm talking about would be (in my opinion, anyway).
This is basically an opportunistic, distributed filesystem that's designed to work with high latency links. Only the tracker can write to its nodes, but anyone can read.
Being able to specify "use 50% of storage to cache low-traffic things, and the other 50% for high-traffic things" would be amazing.
this already exists with addons for Kodi
For your idea, once all the local storage everywhere is filled up with evenly distributed redundant copies, and then a new file is added, would peers arbitrarly choose other files to delete in order to make room for the new file?
For example, Perfect Dark, Winny, and to less extent, Share (which is more similar to eDonkey/eMule).
https://en.wikipedia.org/wiki/Freenet
Regarding of how easy to make something like this network, I'd wager it's pretty hard. There will be a lot of questions, even while establishing the happy path, for example how you manage the updates, especially when you update the protocol, not just the software, and how you effectively manage the volume of search requests, how you distribute the files etc.
And then there's the abuse the network will inevitably get. How you handle spammers, CSAM, malware, ISPs that throttle/block you, the legal risk you put your clients up to, etc. Nice big can of worms. To begin opening it, I suggest a reading through Wikipedia's Peer to peer file sharing article, and especially the File sharing modal on the right, which nicely captures the ideas that have been tried so far.
Only those entities would be able to push data to you, nobody else, so if you trust the providers you specified, you should be good.
Further this lead me to distributed data stores: https://en.wikipedia.org/wiki/Distributed_data_store#Peer_ne...
[1]: https://git-annex.branchable.com/design/iabackup/ [2]: https://wiki.archiveteam.org/index.php/INTERNETARCHIVE.BAK/g...
The list idea could be extended to nested lists (stavros recommends Internet Archive) for discoverability and composition.
If you go with v2 or hybrid torrents from the beginning you could deduplicate and cross seed files from different collections.
The lists could also be modified to have torrents to exclude, possibly using some salt + rehash idea to make it hard to reverse into a list of e.g. CSAM you don't want to publish as is.
Feels like a neat project that could interoperate nicely with existing torrents.
I believe going public inevitably leads to needing to follow growth, which then eventually corrupts virtuous missions.
I'll give you a hint: https://books.google.com/?hl=en
> Search the world's most comprehensive index of full-text books.
It hasn’t been killed but it is clearly a zombie.
In the years the lawsuits were going on, nearly everyone left the project. And then the lawyers have put in so many red lines that it's nearly impossible to make any changes to it.
Given Google's behaviour since the end of their period of true innovative excellence, I don't cut them much slack.
Edit: actually, I think this would make a good submission. Looks like it hasn't been posted since 2017.
[0]: https://www.theatlantic.com/technology/archive/2017/04/the-t...
> People have been trying to build a library like this for ages—to do so, they’ve said, would be to erect one of the great humanitarian artifacts of all time—and here we’ve done the work to make it real and we were about to give it to the world and now, instead, it’s 50 or 60 petabytes on disk, and the only people who can see it are half a dozen engineers on the project who happen to have access because they’re the ones responsible for locking it up.
Could have been great but it wasn't perfect for everyone so it was scraped because surely Congress would take up this noble quest...
If only it still worked well enough to use. I use to use it every day, but how the mighty have fallen.
>> Search the world's most comprehensive index of full-text books.
I mean, the domain resolves, but that doesn't mean the product exists. You can run searches, but you're not allowed to see the results.
I search, see the results, and then click on a result to view the applicable contents of a particular book. In the contents that get displayed, the search term I had entered is highlighted.
a) sometimes the snippet returned doesn't even contain the search term
b) you'll only get the snippets for the first three results, and that's that.
There might be even cases where the book is indexed, but the publisher has disabled even snippets, but I'm not sure about that. But a very limited snippets-only search instead of a full preview is certainly a thing.
> Ambar is an open-source document search engine with automated crawling, OCR, tagging and instant full-text search
> *Easily deploy Ambar with a single docker-compose file; *Perform Google-like search through your documents and contents of your images; *Tag your documents; *Use a simple REST API to integrate Ambar into your workflow
The same was true of the original GOOG 411, which provided a free service, but was really put in place to train up their voice recognition projects.
This is a long running strategy of Google, and it's a shrewd one. The main thing is not to mistake it for a public good. It is an act of privatization.
I think Google initially wanted to augment the web results with a large book collection to get "all the world information and make it searchable", same with Google News.
... pull the other one.
Okay, your statement could be true depending on what you mean by "large". But what makes you think that companies like Google haven't been working on language models without releasing them and/or without discussing them publicly? There's an advantage to be had by keeping corporate secrets.
They offer a clearnet and a hidden service .onion incase you don’t want ISPs blocking access to it.
I don't search for websites based on their titles...
I believe in free access to education, but charging for these books they have no rights to is a whole other thing.
It's a figure of speech, not a comment on what they do.
That being said, 10 downloads/day feels a bit restrictive to me. I'd get if it was 100, or 50, heck, maybe even 20. I mean, I don't appreciate that it's not mirrorable in the first place, but maybe they cannot afford it, I don't know... But 10 feels less than somebody researching a new topic might need to access in a day, even if he won't read them all immediately.
…That being said as well, it has some really nice UI. I wish somebody did it for Libgen.
And if we also count papers, which this site provides too — easily.
Without sarcasm, I don't think the bandwidth bill cares that you find it restrictive. That even more than a handful are free every day for every account is honestly a lot, since that means virtually nobody will need to contribute to the costs they're collectively incurring. And if you're unable to pay, you can still skim a few dozen books (making two or three accounts isn't that hard to do by hand) every day, and go back to any you've already downloaded previously too. And offer them to friends to offload the server.
8 times outta 10, the book would be there.
Well they are providing a great public service and their system takes money to maintain. It's just a small fee.
I realize that neither source will satisfy many of the people on HN, simply because there is a need for current technical books.
Total compliance with the regulatory capture of publishing companies? Full support of the landgrab claims of the Disney corporation et al?
You don't have to be an anarchist to look at the status quo and think some amount of civil disobedience is the correct, proper and right thing to do. Also believing that a claim, fully Disney supported, that such an amount of civil disobedience is somehow ethically evil is, in fact, somewhat disreputable. As disreputable as the insanely high journal subscription fees for taxpayer funded research, for example.
The publishing companies chose this path willingly and with prejudice for their profit turning the relevant law against the people. Are they reputable given they did so? It's hardly an outlying position around here to think they really aren't anything of the sort. Refusing to accept that on mass, until appropriate reform is supported and enacted could be considered quite worthwhile.
You may of course, disagree.
Of course, access to publicly funded research is another issue altogether (and one that should consider patents as well as copyrights).
The question becomes, how many of those targets of "civil disobedience" have earned that privilege and how many are simply collateral damage? Also, ideally, civil disobedience should be aimed towards changing laws rather than people simply taking what they want.
I said "some amount" and you don't disagree. We're not discussing what that precise amount is or how it should be targeted, that's a different discussion and can be framed in multiple, competing ways.
Feel free to lay out what you think is the correct amount and correctly targeted but even if you are wildly wrong in your thoughts, your being wrong doesn't make anyone else with differing thoughts about that amount and targeting disreputable, even if they are for some other reason.
There is a legitimate argument that the law should always be followed and any law-breaking is inherently disreputable. It's not my argument although I acknowledge it and isn't fashionable nowadays given how that must be applied to, for example, the civil rights movement & Dr King.
7TB is even a commodity disk these days. And it's a lot less than the torrent of scientific papers that floated around some time ago (that was ~18TB IIRC).
What I also found was that many of the images in the epubs were already unuseable and nothing like their counter parts in phsyical books.
Kiwix .ZIM file format is a good example. The entire Gutenberg Library is a single ~65 Gb file, and you can read any book from it without unpacking anything.
[1] https://www.timeshighereducation.com/opinion/2048-informatio...
Since transient, ethereal meme culture is also basically emergent culture now, it's difficult not to also foresee a greater cultural divide in such a case. This is saying nothing of live data tools as well, even weather data...
It would be mostly unaffected. As another commenter (Michael) says, communication, news and collaboration are distinct from storage as needs/functions.
What you say here is fascinating:
> Since transient, ethereal meme culture is also basically emergent culture now, it's difficult not to also foresee a greater cultural divide in such a case.
There have always been bookish or gossipy people. Jane Austen and Thomas Hardy both make note of that in stories about English culture. The balance of those qualities may have changed in the internet age. Perhaps the degree of reach via one-to-many communications has amplified the ephemeral gossip side. The scholarly life less so. Though it's enhanced by repositories like Gutenberg, Internet Archive and SciHub and the like, a "reader" (to use Bill Hicks's take) can still only process one media at a time. But as Ted Nelson pointed out, if you have all of the worlds writing at your fingertips, all hyperlinked and with awesome semantic search tools, reading becomes a quite different non-linear experience. That's something centralised walled gardens subtracted from the WWW as its 1990s conceit.
That 90's vision of a multi-modal, multi-media "internet community" in which people read common news, converse and reference together hasn't really survived "Social Media". But then it was always a weal approximation to something like a group seminar in the university library rather than the town square or local pub.
Books are naturally immutable, and could be structured into sub-categories whilst enjoying the benefits of deduplication.
But we were talking about filecoin raking something down from ipfs. They can't do it.
Being coined as the Napster on Blockchain is the worst possible PR one can think of in their situation.
I want a slightly different system, which I've posted about here: https://news.ycombinator.com/item?id=31972252
When you pin files, you announce to the whole world which files you're hosting. When you download files, you announce to the whole world what you're looking to download.
Torrents are simpler and more efficient for distribution, but IPFS is better for accessing individual files.
Somewhere around last year they had datacenters burning down due to a natural disaster, but I'm hoping they can recover. It's an amazing project.
Enabled users: 5629
Active today: 986
Active this week: 2542
Active this month: 4228
Torrents: 566598
Total Size: 23.40 TiB
Retail Torrents: 421313
Creators: 401395
Seeders: 2421540
Leechers: 232
Snatches: 16523563
Transferred: 398.22 TiB
And those are only books. Library of Alexandria is already here.The only problem is that there is so many books and so little time :( , which might be a far bigger problem than book accessibility. To find the time to read them.
There is at least 1/8 of books that I would love to read. But I don't have the time to actually pull it off, even if I am not doing any social networking etc., but the amount is really huge. Maybe, for a startup, someone could index all the content, not only share it.
They couldn't even spring for a Let's Encrypt cert?
When the prosecutor is looking through your internet records and they see 50 wikipedia hits in some relevant time period, they're going to be upset that https exists.
If somebody MITMs it, they can serve you anything they want.
Great. More books!
No really, I don't understand this argument. A static site served by plain http is perfectly appropriate. It's like a poster hanging on the wall for all to see. Of course people can paint over it, but it doesn't really matter.
I agree that it would be an extreme measure.
If it was a quick hack, then that makes it even worse. Difficulty is a mitigating factor.
https://blog.mozilla.org/security/2021/01/07/encrypted-clien...
https://blog.cloudflare.com/handshake-encryption-endgame-an-...
The "anonymity set" refers to the number of possible domains using a single IP address. The existence of that term implies that some IP addresses must have a number of domains associated with them, greater than 1. With these IP addresses, one cannot determine the domain name, the one that the www user sent, from a PTR query alone. Even prior to the introduction of SNI to TLS, when the only way to offer HTTPS was by using a dedicated IP address, discovering the contents of the encrypted Host header via reverse DNS was neither easy nor reliable.
If there are still people reading HN who believe that reverse DNS is reliable and makes plaintext SNI and ECH moot, and are going to comment as such in the future, I would be happy to post the results of an experiment where I take the DNS data for all the domains currently submitted to HN, i.e., a list of IP addresses found in the A records for these names, and do a PTR on each one. We can look at whether "most domains" are identifiable through PTR records.
Also remember the question is not whether ECH protects 100% from someone discovering what domain name the user sent. It does not. The question is whether ECH makes it more difficult to discover than simply sniffing plaintext SNI on the wire, which, of course, is even easier and more reliable than reverse DNS.
The 7TB is via torrent, not via HTTPS. No rDNS needed.
If you use DoH, yes it does. Unless I'm mistaken. They only know the IP address of the remote server.
It's the internet. Everyone can scrape links and measure/correlate which assets were on them to correlate likely visited websites.
Especially if every web page these days is pretty unique in terms of what kind of assets (network streams) with what kind of byte size were loaded at which point in the document loading timeline.
Now include the TLS fingerprint of your web browser and well, privacy went to shit.
HTTP needs an upgrade with scattering and rerouting on the fly, otherwise these deanonymization techniques can never be fixed.
Isn't that e.g. TOR's job? Doesn't belong in HTTP.
If the user requests the page from Internet Archive, Common Crawl or even Google Cache, how does the ISP know what the user requested. (NB. Neither IA nor Google Cache require sending SNI,^1 so the ISP may only see IP addresses).
With IA, the IP address alone does not reveal which IA site or page the user is requesting. There is more to IA than only Wayback Machine.
With Common Crawl, the user can send the Cloudfront domain name instead of a commoncrawl.org domain. Are all ISPs going to know that this is Common Crawl. Even if they expend the effort to learn, what benefit is achieved.
With Google Cache, the IP address alone does not reveal which Google site the user is accessing. Needless to say, there are many, many domains using these IP addresses.
There is nothing that requires any web user to retrieve web pages from a given host. The page may be mirrored at a number of hosts. Some of those hosts might offer HTTPS, support TLS1.3 and not require plaintext SNI/offer encrypted ClientHello.
Even assuming an ISP can determine what domain name a customer is sending in a Host header or ClientHello packet, it would still be necessary to subpoena the archive/CDN/cache to figure out precisely what pages were being requested.
1. The same party is controlling all the server certificates. IA controls the certificates for all IA domains, Amazon (issues and) controls all the certificates for Cloudfront customers and Google controls all the certificates for Google domains. Perhaps there are web users commenting on HN who believe that ingress/egress traffic for site saved/hosted/cached at an archive/CDN/cache is somehow private as against the company running the archive/CDN/cache in a meaningful way. I am not one of them.
As for the question of an ISP modifying the contents of web pages, this is an issue that could be addressed contractually in a subscriber agreement. It stands to reason that if this was a serious issue and not merely a hypothetical one raised by nerds debating the merits of TLS then it would be addressed in such agreements.
As for the "injection of advertising" issue as a argument in favour of the way TLS^2 is being administered on the web, IMO this is a bit silly since (a) it is trivial to filter such advertising (e.g., Javascript in the examples I saw) out out of the page and/or block it from running/connecting/loading and (b) the amount of "tech" company-mediated advertising that web users endure in spite of using TLS is enormous. More likely than being seen as a threat to web users, the injection of advertising by ISPs was seen as a threat to the advertising revenue of "tech" companies. The later are responsible for facilitating the injection of advertising (by their customers, not their competitors, i.e., ISPs), not preventing it.
2. By "TLS administration" I do not mean encryption as a concept nor certificates as a concept. I mean TLS administration measures designed to support "tech" companies first and web users second, if at all. A system where the questions of "threat model" and "trust" are both decided by "tech" companies not users.
TLS itself is only useful if you also rely on DNS over HTTPS/TLS. Well, setting the issues with TLS 1.2 and earlier aside.
There's no reason to ignore a good solution just because it's not 100% perfect.
mine doesn't.
I'm not that offended, and torrents are only available via TOR anyway, but I do actually appreciate the sentiment. There's no reason to be not using TLS.
For some reason I'm seeing this mistake more and more lately. https://en.wiktionary.org/wiki/weary vs https://en.wiktionary.org/wiki/wary
Otherwise we end up in lose/loose situation where I see more people use it wrongly than correctly.
Beyond that, google searches are not indicative of current usage, in SMS/messages, emails, etc. Google search spans all webpages, and many text pre-web too.
I wonder if there's a kind of memetic effect happening online, where people who lack confidence in their English spelling ability see somebody make this mistake and somehow think that it's correct, so they switch how they write it.
I have come across similar use of "revert" in private correspondence from old-fashioned lawyers at small high-street law firms in England. As a rule of thumb, if a solicitor writes that they will "revert to you shortly" you may expect:
* They are fairly close to retirement. * They don't know much about laws that have changed since about 1980. * They almost certainly won't get back to you in a reasonable time unless you repeatedly send reminders. They are probably hoping that they will be able to retire before having to invent a meaningful response to your enquiry.
I notice it more and more too.
Probably 100% doesn’t matter in this case.
This has to be a lot of duplicates or bad formats (images). This would be far more useful to people with some curating.
The top ten countries average north of a hundred thousand books per year each[1], so let's say a million books per year globally because I haven't got the time to extract the numbers and sum them all up exactly. It probably wasn't as much fifty years ago, but we're also completely ignoring the Internet which is way bigger because of user-contributed content (not to mention things like newspapers, meeting notes, etc.), so let's say this was the case for the past fifty years. That's fifty million books. Average book has 85k words[2], and a word is like 5 characters on average. Not all countries write in English, especially some big ones like India and China, so we could probably double it on a global scale but let's go for a conservative 7.5 bytes per word and add a byte for the space (I'm ignoring other punctuation). That comes out to 8.5×85e3×1e6 which is less than a gigabyte and uncompressed. Decent compression iirc makes it a fifth of the size, so 145MB compressed.
I've still got to be an order of magnitude off because we write a whole lot more than just books (and even if we were just looking at books: there are also book revisions, drafts, etc. I'm just counting the published words), but that's still less than expected.
In conclusion, it comes out to about 1/50'000th of 7TB compressed, which I would say lends some credence to the claim—even if I feel like I must have made a mistake somewhere because 145MB for ~all books of the past 50 years from all countries seems quite little.
[1] https://en.wikipedia.org/wiki/Books_published_per_country_pe...
[2] a few sources on ddg, e.g. https://www.tckpublishing.com/how-long-should-a-book-be/
Presently, only the first four of several dozens of parts are available.
You don't need book to be smart... you need a brain :)