One good self-hosted OCR is PaddleOCR, https://github.com/PaddlePaddle/PaddleOCR
Beats everything else, truly international and multi-lingual, including Chinese (as it is made in China)
Beats everything else, truly international and multi-lingual, including Chinese (as it is made in China)
Their PaddleLayout models are also miles ahead compared to LayoutParser or TableTransformers in both inference speed and output quality
The reason I use it is to test whether it's just analyzing letter-by-letter (even if they claim it does more) or if it's actually scanning the letter/word in its context. If it's letter-by-letter, I get hilariously awful results.
Sure, it got things wrong. But it also figured out some things even I couldn't decipher.