I think perhaps the developers took a wrong turn when they started trying to improve OCR with language models rather than font models. Humans can accurately transcribe a printed text without knowing the language. They do that with a mental model of font metrics and so on. In fact, a human transcribes more accurately (though more slowly) when they don't know the language because they don't erroneously "autocorrect".
I wonder if Google has in-house OCR that works much better than commercially available OCR. Google has OCR-ed at least 25 million books. You can't download the complete texts, but you can see snippets. Perhaps someone would like to publish a paper assessing the quality of Google's OCR and comparing it with commercial software. (Probably someone has done that already; I'm just bad at finding papers.)