If maintainers are seeing this: any plans to publish non-English books? Is there room to collaborate on adding support for it?
If maintainers are seeing this: any plans to publish non-English books? Is there room to collaborate on adding support for it?
One fact of reality is that Standard Ebooks' tooling is English-based. Everything from the pages/xhtml it generates to its typography tools to its style guide expects English. For example, Spanish uses — and «» instead of English's “” and ‘’ for dialogue so you can imagine how SE's punctuation tooling heuristics are going to differ here.
You'd also have to come up with a new set of standards for another language. What sort of correction is a fair modernization and which would be unfair editorialization? SE itself already makes controversial decisions here for English like "to-day" -> "today".
It would be a large undertaking to parameterize some sort of LANG=EN setting and imo not worth it. I think the only route for that sort of thing to happen is if 2+ "forks" get to Standard Ebooks' quality and then decide to work together years down the road.
Also, legal clarity is sometimes an issue in other countries where an English translation written in the USA of that same work is clearly in the public domain. There are works where the original non-English content, despite being long translated into English, are still owned by the estate, but the English translation is liberated.
Something I realized was just how many books exist only as scans. Transcription isn't very fun, but this makes it rewarding. For example, I have some early Spanish sci-fi books I've transcribed into epubs that you cannot find outside of dirty scans.
https://standardebooks.org/contribute/accepted-ebooks
> Types of ebooks we don’t accept
> Non-English-language books. Translations to English are, of course, OK.
1. Discover / browse / search.
2. Expect high quality text & typography.
3. Be part of a high-volume community that archives and refines texts year after year.
"Then build your own similar site with high-quality titles in French!" , I hear some say. Sure, but I see tons of excellent marketing & infrastructure already done by Standard Ebooks, which could be reused for non-English books!
Said differently, mid/long term I see the greater good being achieved by supporting non-English titles in Standard Ebooks itself, not in one satellite mimic site per language, of varying maintenance quality. To compare with the domain of online encyclopedias, these are the same reasons to have en.wikipedia.org and fr.wikipedia.org maintained under one (Wikipedia) umbrella sharing infrastructure, not wikipedia.org and frenchwikipedia.org maintained by entirely different teams.
[1] https://groups.google.com/u/1/g/standardebooks/c/JdVpCm3ckGg...
[2] https://groups.google.com/g/standardebooks/c/osOEfs5HdLo/m/2...
Agreed, it's worth bringing up again on the mailing list; will do this week and post here a link to my message.
This isn't the fault of Standard Ebooks, but it's an elephant in the room of public domain texts in general. In many ways I'd rather see effort put into crowdsourced or volunteer modern translation than into improved copyediting.
It's especially problematic if you consider that the original text in the foreign language isn't available as well, as you're noting.
Mistakes can/will be done, and fixed afterwards. It's the beauty of our online worlds.
All of the work in polishing an epub is the chore of transcription and then nitty gritty details like correctly tagging things like roman numerals and embedded poems and following some sort of standardization guide.
The thing that Standard Ebooks does, aside from being decisive over its standards, is then require the epub to go through a review process by its creator who is a domain expert in the craft. This expert bottleneck is a big reason why SE book quality is so reliable. Accepting other languages drastically changes and perhaps even relaxes this bottleneck and changes the whole organization.
In another comment, I think you suggest that it might be time to try to persuade him to accept non-English books:
> Agreed, it's worth bringing up again on the mailing list; will do this week and post here a link to my message.
But it's not really up to persuasion. Because it takes more than an idea to expand the accepted languages. It requires at least one reliable expert in that language who can stand up a completely new set of tools and standards and then steward that project to fruition. And that's such a big undertaking that it's really a whole new project, not just yet another egg under SE's wings, but a whole new chicken coop.
Suggesting that Standard Ebooks move to support other languages is 0.0000001% of the work towards that goal. People do that all the time, then claim "okay, I'll fork the project", and then fizzle out.
Another example is that whole swathes of code in https://github.com/standardebooks/tools, SE's core workflow, become useless once you're targeting something other than English. By browsing their style guide and that repo, you'll realize that SE's value really is its focus on English. It's more of a suite of English tools than it is a epub editing kit, as the latter is the easy part.
I've been monkeying around open source long enough to know veeeery well that "Suggesting [...] is 0.0000001% of the work towards that goal" , and I've been culprit of that myself :) .
Still, I like SE a lot and might be interested in doing the work. So, I'll make my point to the ML (expanding on A. points I brought here, and B. contradictions that you and other commenters wrote, thanks), asking if folks are convinced by the vision, and asking for technical advice to build the incremental path. Then if there's agreement, maybe I, or someone else, will commit to it.
It's also the only avenue that makes sense. As I said in a sibling comment, I've been working on a Spanish-language version of SE and I have some good ideas of how the tool chain could be parameterized for multi-language support, but it also feels like a pointless cherry on top (and a complication of an already non-trivial workflow) to bring everything under one umbrella. And it would require a loss of control for SE's creator.
I recommend epub'ing a few book scans yourself in your preferred language and starting a project around that. You'll find other people who have at least started a fork themselves that you might be able to round up under one org. I personally have aimed for 10 ~finished books and a domain name that hosts my own style guide + tutorial to prove that I'm serious about it (to myself) before I publish my efforts. And you need as much to appeal to any would-be contributors.
I think the only plausible place to be in 5 years is for a couple major sister projects to reach maturation and then form a sort of ring of "Check out lettres-libérées.org for a similar project for French works."
Thinking about it! Thanks for the advice.
In the past people have expressed interest in forking the toolset for other languages, which is totally fine. But I don't think I've seen any of those attempts come to fruition yet.
In another comment of this sub-thread ( https://news.ycombinator.com/item?id=25140605 ) I wrote:
> "Still, I like SE a lot and might be interested in doing the work. So, I'll make my point to the ML (expanding on A. points I brought here, and B. contradictions that commenters brought), asking if folks are convinced by the vision, and asking for technical advice to build the incremental path. Then if there's agreement, maybe I, or someone else, will commit to it."
Do you think it remains worthwhile that I bring the discussion, or is it 100% nailed among current maintainers that the "incremental path" I'm hoping for doesn't exist and "just fork / do your own thing for your language and register your domain" is already the consensus?
(which are sourced from Gutenberg catalog, but with dynamic navigation and search)