HNHacker News
TopNewBestAskShowJobs

weijiacheng

51 karma · joined April 28, 2021

submissionscomments
weijiacheng··on Standard Ebooks
One of the funny things about Bible translations is that more modern translations are based on older manuscripts than older translations, due to advances in archeology. SE can't carry any translations that incorporate the insights of the Dead Sea Scrolls, and having access to some of the oldest Hebrew manuscripts is a pretty big deal when it comes to translating the Tanakh.

It's true, modern versions of War and Peace can't be hosted at SE, but those modern versions generally don't reflect revolutionary leaps in archeology :)

weijiacheng··on Standard Ebooks
Yes, bookshops will sell one version of the Bible to Catholics, another to Protestants, another to fundamentalists, another to progressives, etc. :)

In contrast, part of the SE editorial philosophy is that it tries to host the best (based on academic scholarship, translation quality, academic acclaim, etc.) version of each text available in the public domain, which excludes that "something for everyone" sort of play available to a commercial bookstore. You could rightly argue that this is losing something (it's good to have multiple translations to compare if you're reading a text for critical purposes), but the SE editorial philosophy avoids a certain amount of confusion and clutter for the general reader. So there's a deliberate (you could call it "arbitrary" in some sense, if you wish) tradeoff being made here.

weijiacheng··on Standard Ebooks
I could see how that would be helpful, but at least for my use case I'm more interested in seeing how LLMs integrated with computer vision can speed up transcriptions. Since a thorough proofread by a human is already baked into the SE production process (and is indeed one of the major selling points), having more automated tools to aid proofreading is nice but doesn't do anything fundamentally different, from my point of view. Whereas if LLMs can be leveraged for transcription SE producers no longer need to depend on external projects like Project Gutenberg or Wikisource to produce texts (which can take months) or transcribe texts from OCR results by hand (very tedious and error-prone--believe me, I'm speaking from experience!). It would drastically open up the range of possible books someone could reasonably produce (in a timely fashion) for SE.
weijiacheng··on Standard Ebooks
The site actually hosts several "religious books" (try filtering by the "Spirituality" tag -- I've even produced several books on religious topics myself for SE). What it doesn't host are "Religious texts from modern world religions" (what some might call "scriptures," e.g. the Bible or the Quran) which is a much narrower category than "religious books."

As a religious person myself, I actually think this policy is very sensible. Most (nearly all?) religious texts of major world religions were originally written in languages other than English, and so if SE were to try to host those texts the site would have to make an editorial call about which translations of those texts are the "best." That quickly enters very murky theological territory, where one side of a given religion might push for one particular translation, whereas another side would push for another translation.

To give the Bible as an example, Catholics and Orthodox Christians include the deuterocanonical books (e.g. Tobit, Judith, Sirach) in their canons whereas Protestants exclude these. Would the SE version of the Bible include these? Some American fundamentalist Christians claim that the King James Version is the only valid English translation of the Bible, whereas the Revised Version (also available in the public domain) is based on more reliable Greek manuscripts. But some conservative Christians reject the Revised Version and its descendants based on certain theological premises...

Do you catch my drift? IMHO it's very sensible for SE to avoid these sorts of debates entirely by sticking to books where you could argue (with some degree of handwaving) that there really is a "best version" :)

weijiacheng··on Standard Ebooks
In addition to what Alex has said, as an SE contributor I do try to submit errata to Project Gutenberg where I can find the time and energy. Part of the problem, though, is that PG's errata process (https://www.gutenberg.org/help/errata.html) is quite cumbersome since you have to write an email to their errata team with each individual error. That's a real hassle to try to keep track of and submit. Ideally, if PG had something like a pull request system, I would just be able to find those errors in their code and submit the changes directly, but unfortunately they don't have that, so far as I am aware.

That is one major advantage SE has, I think, which is that we do allow people to make pull requests against any of our ebook repositories and any PRs that get merged are automatically deployed to the site. This makes it much, much easier for tech-savvy people to submit proofreading corrections!

weijiacheng··on Standard Ebooks
I am one of the SE editors/regular contributors and I did play around with this a bit for a poetry collection: https://groups.google.com/g/standardebooks/c/IUvGLmvZrmM/m/s...

I'm sure someone sufficiently determined and good at prompt engineering, and integrating LLMs into a larger toolset, could come up with something even better. I'm personally very skeptical of LLMs as a technology, but even I have to admit that this was a pretty ideal and unobjectionable use of LLMs.

That being said, though it was a fun experiment, I later found that it was easier (and less wasteful of natural resources) to just do the same thing with a bit of custom markup and a search and replace script.