PG does great work and we rely on them almost exclusively for transcriptions. But we're two friends working towards to different goals.
PG does great work and we rely on them almost exclusively for transcriptions. But we're two friends working towards to different goals.
Until I got to this part of the comment I was thinking "Yay, an alternative to PG's godawful OCR transcriptions". Why would you reuse the worst part of Project Gutenberg?
When I did "The Valley of Fear" as my first project, the PG text was used as the base, but if I encountered any kind of ambiguity in the text, I consulted at least a half-dozen other versions of the text via Google Books scans for agreement.
The team is also very particular about only using editions that have entered into the public domain. So if the first edition of a book just entered public domain, you must make sure that what you have produced only uses text from the first edition, and that you haven't inadvertently used a later edition as a base that may have included subsequent editorial changes.