HNHacker News
TopNewBestAskShowJobs

tangledhelix

23 karma · joined May 20, 2013

submissionscomments
tangledhelix··on Project Gutenberg – keeps getting better
Since the books are available on the site as text and HTML the search engines index them already for you. Try searching for the below; it should take you to the book you expect as the first result:

site:gutenberg.org "it was the best of times"

tangledhelix··on Project Gutenberg – keeps getting better
One author remains blocked in Germany (but only for a couple more years)...
tangledhelix··on Project Gutenberg – keeps getting better
This is covered in the FAQ - https://www.gutenberg.org/help/faq.html#why-is-project-guten...

And as another person noted, the vast majority of books have HTML, EPUB, Mobi formats. We are also looking at both KEPUB (Kobo) and PDF which will probably come in the future.

tangledhelix··on Project Gutenberg – keeps getting better
As another commenter said PG is almost all books from 95+ years in the past due to copyright law in the US. We partner with a sister organization, the World Library Foundation, who have a self-publishing portal for modern works by authors who wish to put their own work in the public domain. You might want to look there for more modern material. https://self.gutenberg.org
tangledhelix··on Project Gutenberg – keeps getting better
Cloudflare sells that as a product, they call it Labyrinth IIRC.
tangledhelix··on Project Gutenberg – keeps getting better
See what I wrote above (and let me say I am talking about Project Gutenberg and Distributed Proofreaders here, I am one of the admins on both). A large amount of the hassle traffic we've seen is as I wrote above, the IPs come from everywhere and in many cases, each IP makes a single request and doesn't come back. They change user-agent dynamically, etc, to masquerade as regular traffic. They come from residential, cloud/hyperscale, corporate, educational, government, all the networks, on every continent. This is many thousands of "open a ticket with someone" events per hour territory. It's as difficult to fight as DDoS itself for the same reasons (presumably the harvesting parties know that and that's exactly why this approach is used).

Others online have been writing about their own experience with the same stuff; it's not unique to PG at all, it's everywhere. Talk to anyone that runs a web server and they'll have these stories...

tangledhelix··on Project Gutenberg – keeps getting better
One could argue that this falls into the previous poster's thought about "the little differences to modern English are part of the charm" ...
tangledhelix··on Project Gutenberg – keeps getting better
OCR has improved a lot since then, but OCR is just step 1 of reading in text. They make a lot of errors (even now, especially on old worn out paper pages) and even if they didn't, one has to format the book, deal with footnotes, sidenotes, illustrations, etc. DP is very active, we will welcome you back with open arms :)
tangledhelix··on Project Gutenberg – keeps getting better
There are many books available as audio, some are human-read, some were automated. You can see lists here:

human-read: https://www.gutenberg.org/browse/categories/1

computer-generated: https://www.gutenberg.org/browse/categories/2

IIRC many of the human-generated ones come from LibriVox, many of the computer-generated ones came from a collaboration with Microsoft.

tangledhelix··on Project Gutenberg – keeps getting better
Worse than that - even if they would take action, you can't possibly orchestrate filing all of the complaints. It's a drown-in-quicksand problem, you can't fight quicksand one grain at a time.
tangledhelix··on Project Gutenberg – keeps getting better
The ebook editions are very good for this. Most of the e-reader software provides all the amenities (bookmarks, highlighting, notes, control of margins, etc).
tangledhelix··on Project Gutenberg – keeps getting better
The Alfred Döblin books are still blocked in Germany (for a couple more years).
tangledhelix··on Project Gutenberg – keeps getting better
wouldn't help, much of the traffic we've observed look closer to ddos patterns - IPs from all over the world, many different networks, each IP makes one request only, doesn't come back. highly distributed, no form of blocking would be effective except maybe captcha or proof of work.