Standard Ebooks
standardebooks.org
standardebooks.org
Of interest might be my blog post on how SE runs on a small VPS using classic web tech: https://alexcabal.com/posts/standard-ebooks-and-classic-web-...
(This post is slightly out of date as there is a database now; but it's used for managing Patrons - and soon a cover art listing and approval system - not for serving the actual ebooks, which are still served as described in the post.)
Our volunteers have spent the last few months preparing a few notable books published in 1928 to be released today, Public Domain Day. Those are the top 5 books in the ebook list, starting with The Mystery of the Blue Train. Check them out!
We welcome new contributors if you'd like to work on producing a new ebook. In the next week we'll also have a brand new cover art database launched, so if you'd rather help by cataloguing new cover art for future ebooks, get in touch at our mailing list!
Have you considered making books sortable by popularity? It might be more approachable for new users if they see books they recognize at the top.
Edit: oh and Lambda for a total of 2 server functions
I remember when Manybooks used to be what you want. But quality dropped precipitously with self-published new novels, I suspect some money is changing hands somewhere.
What happened to Manybooks? Does Standard Books have a plan for avoiding that?
At SE we focus exclusively on US public domain titles; that's one of the major philosophical points of the project. The other major point is a high quality standard, so it's in our best interest to keep pursuing that. SE became known due to its quality standard, not because it's more free ebooks. Therefore if we strayed from those points then we'd be just another free ebook site, of which there are no shortage.
Quality is also why we reject self-published books that have been dedicated to the public domain, as those are typically low-quality content to begin with. (Though I wouldn't call every single book we host "high quality content" in the sense that each one is up there with Shakespeare. But books that have survived a hundred years tend to have survived because they're not slush.)
Every other book I read now is by an author with NO rating, I have read six this year, none were memorable, or my cup of tea I will give you that, but two of the four- or five-star offerings on Amazon were just as bad. As they say, if you don't open an oyster, you will never find a pearl.
That's not the niche SE has chosen to target. You can't expect them to serve every possible use case.
There's been some mailing list chatter lately on how to best format PDF editions, but that's not being pursued on a project level.
A hurdle for this though, is that building a good website for the Kindle browser is a pain, as the browser's support for various html/css/js features and standards is all over the place, with no debugging tools available.
I say the same thing in every ebook thread: On a purely technical level Kindle is a terrible ereader designed by people who seem to hate books. Buy almost anything else.
Weirdly, about half of them have developed a problem after about 5-7 years, whereby they intermittently stop charging. Replacing the battery doesn't seem to fix it. Might be a problem with the soldering of the USB connector to the PCB?
As a bonus they are Linux based, and you can do fun things like replace internal SD cards with bigger ones, login using telnet and install new applications.
They're also quite a nice embedded ARM Linux machine for a lot less than I could make one or buy one from elsewhere, but I suspect that isn't the core market for a kindle...
- "Browse free ebooks in the Encyclopædia Britannica’s Gateway to the Great Books set[…]"[1]
- "Browse free ebooks in the Modern Library’s 100 Best Novels set[…]"[2]
- "Browse free ebooks in the Modern Library’s 100 Best Nonfiction set[…]"[3]
(Indeed, the titles are even much longer than that. It feels SEO-ish; not sure why that would be a priority for a free culture project like Standard Ebooks, especially give the momentum and cachet it already has.)
Collections should also have placeholders for unavailable titles. For example, currently the "Utopian Trilogy" collection[4] contains exactly one item, in spite of the true size of the set it actually belongs to. When an item is not available because of copyright, that (along with the year in which SE will first be allowed to make its own edition available) should be made clear. Where it's unavailable because no one has yet proofed the text for an SE edition, a clear call to action can be made.
And it's seemingly minor, but on the subject of editions, I wish SE followed closer to the print tradition instead of the modern Web millieu and clearly identified its microeditions as exactly that: distinct editions of the same text. (Yes, that means there are possibly dozens (or hundreds?) of different editions, given that errors can be found after the fact and the SE house style may even change, necessitating updates. No, that's not a problem.)
1. <https://standardebooks.org/collections/encyclopaedia-britann...>
2. <https://standardebooks.org/collections/modern-librarys-100-b...>
3. <https://standardebooks.org/collections/modern-librarys-100-b...>
That looks like it might be search engine optimization.
As a religious person myself, I actually think this policy is very sensible. Most (nearly all?) religious texts of major world religions were originally written in languages other than English, and so if SE were to try to host those texts the site would have to make an editorial call about which translations of those texts are the "best." That quickly enters very murky theological territory, where one side of a given religion might push for one particular translation, whereas another side would push for another translation.
To give the Bible as an example, Catholics and Orthodox Christians include the deuterocanonical books (e.g. Tobit, Judith, Sirach) in their canons whereas Protestants exclude these. Would the SE version of the Bible include these? Some American fundamentalist Christians claim that the King James Version is the only valid English translation of the Bible, whereas the Revised Version (also available in the public domain) is based on more reliable Greek manuscripts. But some conservative Christians reject the Revised Version and its descendants based on certain theological premises...
Do you catch my drift? IMHO it's very sensible for SE to avoid these sorts of debates entirely by sticking to books where you could argue (with some degree of handwaving) that there really is a "best version" :)
In contrast, part of the SE editorial philosophy is that it tries to host the best (based on academic scholarship, translation quality, academic acclaim, etc.) version of each text available in the public domain, which excludes that "something for everyone" sort of play available to a commercial bookstore. You could rightly argue that this is losing something (it's good to have multiple translations to compare if you're reading a text for critical purposes), but the SE editorial philosophy avoids a certain amount of confusion and clutter for the general reader. So there's a deliberate (you could call it "arbitrary" in some sense, if you wish) tradeoff being made here.
So modern versions of e.g. the Bible could not be in Standard Ebooks. So easiest to not carry any translations.
Bookshops have no problem with this as part of the purchase price will go to the copyright owners of the translation.
It's true, modern versions of War and Peace can't be hosted at SE, but those modern versions generally don't reflect revolutionary leaps in archeology :)
There are modern translations that are permissively licensed and are of surprisingly high quality. See the NET Bible as a prime example. It's also the only one I know of with good translation notes that can be had for free.
The site already hosts a number of works that were originally written in languages other than English, and yet it had no problems making an editorial call about which translations of those texts are the "best." The obvious solution would be to just allowing multiple translations of foreign-language books.
Is there a technical reason to disallow multiple translations of the same text? I can see on the "wanted ebooks" page a number of translated titles[0]; so the project does seem to make editorial decisions about which translations to work on. Obviously, where one translation exists, there may be others that have other advantages.
The Didache
Anselm, Cur Deus Homo
Anselm, Proslogion
Augustine, City of God
Augustine, Confessions
Augustine, On Christian Doctrine
William Law, A Serious Call to a Devout and Holy Life
Luther, The Bondage of the Will
Calvin, Institutes of the Christian Religion
Pascal, Pensées
All of these are in public domain.
I understand that this is their fault and not yours, but maybe it could be interesting for you to offer one of these formats now that the Kindle browser is actually usable?
Do they actually check the file content or just the name?
If the latter, just a .AZW alias to the actual .AZW3 might work ...
(I can't test, don't have a Kindle nowadays)
Once the files are in the Kindle, it would probably work out OK.
I'm not complaining - It's just, there's a reason everyone goes for the existing frameworks and it isn't addiction to complexity. Raw PHP code is legendarily insecure and prone to XSS and other issues if you don't do things exactly right.
Nice site, though.
Not any more so than sites with frameworks. I’ve found XSS issues in Java Spring framework built sites that didn’t “do things exactly right”. A framework doesn’t magically fix that.
Thank you for your beautiful project.
For a few years, every January 1st, for public domain day, I have been promoting SE on social media, the thread on Mastodon is the one with the most involvement. https://fosstodon.org/@paulox/111680544393923401
It would be nice to have an SE account on Mastodon that posts about every new book published, since IMHO it's the social network more aligned with the spirit of SE.
If we did, then someone would have to volunteer to run the account, and also the account must be able to delegate posting powers to another user without exposing the account's master password (like Tweetdeck or Facebook are able to do). If that's possible and you're interested in helping, please send me an email!
Kindle will not natively support epub until you can connect it to a USB cable, transfer an epub using a file manager, and it does not get secretly converted.
Almost all my Kobo books are EPUB and work great.
Re XHTML - It looks like the website is being served with a text/html content type. Did you give up on your XHTML experiment? How did it go? Maybe it would help if browsers reported back errors to the server, like how "Content Security Policy" reporting works.
The process and tools are quite nice and it's very rewarding to see your work in ebook form. It takes a _long_ time to proof and re-read a book, but it's surprising how many times you can do this and how differently you need to read to catch errors versus just enjoying the damn book.
The fascinating part of the project is a _strong_ editorial opinion, which IMO makes the project successful. There is a core group of people that upholds the standards for the project, and the resulting consistency of quality of output derives from that. The team clearly cares about the quality, and has demonstrably maintained that over the huge number of releases.
I even went to the archives of the "San Francisco Newletter and California Advertiser" to collect some of Bierce's original work, making it the most complete, and most corrected open-source version of the book. [1] The one previously hosted by Project Gutenburg was quite old and, frankly, quite riddled with transcription errors.
I haven't tried reading the Devil's Dictionary back-to-back since I published it, but I might one day. There's a lot of detail in this work that I never saw until I had it under a microscope.
[0] https://standardebooks.org/ebooks/ambrose-bierce/the-devils-...
[1] https://archive.org/details/san-francisco-newletter-dec-11-1...
[0] https://standardebooks.org/about/what-makes-standard-ebooks-...
So submitting back to PG would be more like replacing a PG edition, instead of updating it; and I doubt the original PG submitter would like it if their hard work was simply replaced by someone else who thought their version was an improvement.
Our volunteers do sometimes submit typos they find back to PG. We don't require that, so some producers do, and others don't.
But yes, I see that you’re practising some editorial oversight and not aiming to faithfully represent the original in all regards, which I gather is more generally Project Gutenberg’s goal; and this would obviously contraindicate upstreaming.
On the other hand, when it comes to more stylistic matters, I tend to wish Project Gutenberg had more consistency. There’s too much gratuitous variation in presentation and ridiculous 256-colour backgrounds. It’s often too obvious much of it is the work of a group of individuals rather than a coherent effort.
I’m curious about the footnote-to-endnote thing, because I’m not sure how the various formats in question handle them all, but in print endnotes are almost always just awful. If anything, I’d be expecting to replace endnotes with footnotes. (Me, I’m partial to sidenotes.)
—⁂—
¹ Hickory dickory dock, three mice ran up the clock; the clock struck one, and has been charged with assault and battery.
Agree on notes in print, side notes (on very-wide editions) are best, then foot, then end of chapter endnotes. Full end-of-work endnotes are awful. Maybe they're better in ebooks, than footnotes, though? E-readers' poor UX for not-even-that-advanced features of books is part of why I barely use them, and practically never for any work that'd have notes of any sort.
In the case of Standard Ebooks, “sound‐alike” changes are allowed (so spelling and capitalization changes are allowed when they make sense). Censorship, and even innocuous grammatical changes, are not. Despite generally appreciating old works in their own context, I find the tradeoff in readability for such a widespread practice to be worth it given how minor SE’s alterations are.
When I see erroneous changes in SE books, I argue to revert, and have generally been successful. In my experience it’s drama‐free, like fixing any other typo.
That is one major advantage SE has, I think, which is that we do allow people to make pull requests against any of our ebook repositories and any PRs that get merged are automatically deployed to the site. This makes it much, much easier for tech-savvy people to submit proofreading corrections!
On the other side of the coin, Standard Ebooks's heavy endorsement/buy-in of GitHub-based workflows are offputting to broader audiences. (It's pretty offputting to me, and I'm not even non-technical; I just recognize it as a sort of Conway's Law + Law of the Hammer sort of thing, and it chafes.) I.e., for others what you describe is far less than "ideal".
Can you share or document how? https://standardebooks.org/contribute suggests that "Technically inclined readers can produce ebooks themselves" but doesn't provide any point of entry to do so other than a link to the GitHub org, and "No technical experience is necessary. Contact the mailing list if you want to help." just links to the Google Group.
Every time I buy one of these public domain books from Amazon, they are invariably shitty, low-quality "printed by Amazon" versions. I miss the time where you could get a high-quality hardcover, but more and more those seem reserved only for the current week's NYT best-seller books.
I feel like these open domain novels published by big publishing houses have the veneer of legitimacy, but projects like the one this thread is about I think could accomplish much more. Especially for authors where the work is translated into English. Plus the cover designs are much cooler.
I will say, the search on their website is kind of slow and could use some work.
https://www.amazon.com/Greatest-Works-Dostoyevsky-Punishment...
Although there's no saying as to whether or not they will have proper spellcheck, TOC, if they are legitimately in the public domain, or even if it's the right book with all the pages. That's where a service like Standard Ebooks is superior to the potluck you get from Amazon.
* NB: whether this is actually the case or not is a separate matter
Regardless, it's great that these works are available in high quality for free.
A physical copy if I wanted one I'd be willing to pay ~$6 for, less if used.
That doesn't make sense. You disagree with what? My expectation that it shouldn't cost much more than $5 for an ebook of something with a 150-year-old plot?
> $10 to be able to have it this second on my Kindle vs. waiting to get the physical copy feels like a good value
That's nice, I guess—for you. But we weren't talking about you, and we weren't talking about instant gratification.
If I'm buying for reasons where instant gratification isn't a factor—and I'd argue that for books, prioritizing instant gratification is even sillier than with e.g. food and drink or streaming TV shows—should I still pay the premium to be able to get something "this second" if my flight isn't for another three weeks and I'd have been perfectly willing to deal with a delay?
If you want the cheaper one buy the cheaper one. If you aren't in a rush to get it, buy the one that gets delivered in a week instead of in 10 seconds. If you want the ebook, buy the ebook. If you want the printed book, buy the printed book.
> it shouldn't cost much more than $5 for an ebook of something with a 150-year-old plot
> for a classic work [...] $3 or $4 sounds about right
That is a direct response to your question to the other commenter about why Penguin's Crime and Punishment for Kindle is a ripoff.
I don’t think italicizing makes the argument more ironclad.
$3 or $4 doesn’t sound about right to me.
To me, the age of something doesn’t necessarily influence its cost. Saying “this book is old so it should be cheap” is something I disagree with.
> $3 or $4 doesn’t sound about right to me
No one asked. You asked, on the other hand, a question about the ripoff comment, and that's the question you got an answer to.
No more attempts at conversational sleights of hand, please.
You cannot say to someone who doesn't like hot dogs, e.g., "What makes hot dogs gross?" and then when someone explains why, respond, "I disagree." That makes no sense. You asked for the information, they gave it to you, so the only possible way that "I disagree" fits at that place in the conversation is if you're saying that you disagree that that's their reason. To respond "I disagree" in the sense that you don't share their taste is to fundamentally change the subject to try to have a different conversation. And it's not interesting, besides.
People have different tastes. That's expected.
Stating that you personally think it is worth the price—especially in this context, as if it's some kind of retort—is just annoying. It isn't illuminating; that you don't share the same opinion is utterly unsurprising and didn't need saying, and it provides no special insight.
If we did offer print books, I think the value-add would be making them extremely ornate, one-of-a-kind editions like Arion Press or Folio Society make, and we'd charge a lot for a copy. But even then I'm still not sure the juice would be worth the squeeze, because that's also been done to death... how many more fancy editions of Dracula or whatever does the world need?
I've been toying with the idea for a while but I think the market is just too saturated, even for premium editions. Maybe the focus should be on reviving more obscure works... not sure.
Standard Ebooks - https://news.ycombinator.com/item?id=32215324 - July 2022 (256 comments)
Free and liberated e-books, carefully produced for the true book lover - https://news.ycombinator.com/item?id=25138534 - Nov 2020 (106 comments)
Standard Ebooks: Free public-domain ebooks, carefully produced - https://news.ycombinator.com/item?id=20594802 - Aug 2019 (129 comments)
Standard Ebooks: Free and liberated ebooks, carefully produced - https://news.ycombinator.com/item?id=14570035 - June 2017 (96 comments)
example:
https://standardebooks.org/ebooks/rudolph-erich-raspe/the-su...
A bit offtopic, but I never understood why .epub is a thing. For instance the linked HTML/XHTML version seems to work just fine (except for the reflow thing.. but I assume that a CSS issue)
.epub seems to be mostly HTML with a few pieces missing. I guess I don't understand why we needed a new format? and not just use a strict HTML subset?
I'd love some strict HTML subset that indicated the file can be used offline. I personally try to make all my webpages so that they can be saved to disk and opened from a single file (though if you embed images/videos this becomes problematic). But I don't have a way to indicate to a reader "Hey you can Ctrl+S this webpage". I'd publish .epub, but the browser won't open them
For example I live in Iceland there is a number of texts that are in the public domain, for example the Edda and Icelandic sagas among others. But since we are very few (approximately 380K people) there is no comparable entity and most likely never will be, so the best and probably only way to get something similar would be to be a part of a larger organization.
If you do decide to start something then SE’s tooling supports a --white-label flag when creating a skeleton, which would at least get the first few productions off the ground.
Feel free to ask on the mailing list if you have any questions, more likely to be picked up there than in a random HN thread :)
Some classiications seem a bit ...nonintuitive. For example, the Autobiography of John Stuart Mill is classified as "very diffcult" whereas "The Tempest" by Shakespeare is classified as "fairly easy".
I would classify it the other way around, but what do I know, I'm a nonnative speaker anyway.
https://en.wikipedia.org/wiki/Flesch%E2%80%93Kincaid_readabi...
I'm sure someone sufficiently determined and good at prompt engineering, and integrating LLMs into a larger toolset, could come up with something even better. I'm personally very skeptical of LLMs as a technology, but even I have to admit that this was a pretty ideal and unobjectionable use of LLMs.
That being said, though it was a fun experiment, I later found that it was easier (and less wasteful of natural resources) to just do the same thing with a bit of custom markup and a search and replace script.
The most natural application of a language model in proofreading is to compute perplexity across the text; if all goes well, errors should be detectable as points of unusually high perplexity. (In principle, this should even be able to spot otherwise undetectable errors like missing words.)
The site is called Modern Serial, and it lets you read books from Standard Ebooks in 10 minutes a day as Substack-style email newsletters.
For now, your best approach would be to take high-quality ebooks like what Standard Ebooks offers, and use text-to-speech software.