Having used Tesseract for OCR for other things, getting the right PSM
helps but it's still rather terrible, especially for sans-serif fonts, which are common in UIs.
Granted there's a lot of ambiguity in sans serif fonts, lower-case "L", vertical bar, and upper-case "i" can even be pixel-identical, but I've seen tesseract turn
Chapter III
into
Chapter |l1
which really surprises me. In fact, for books, I run it through sed to replace vertical bar with upper-case "i" and it significantly improved recognition.