It's not very good. I miss being able to copy/paste from blurry or deformed screenshots of youtube on Windows.
It's not very good. I miss being able to copy/paste from blurry or deformed screenshots of youtube on Windows.
The special sauce - what you need to get a better result - is good, adaptive thresholding (something more advanced that raw naive binary thresholding you get feeding naive color/grayscale images to OCR).
As far as I know, once you get that nailed it doesn't matter that much what OCR you use - as long as it's available and supports your target language.
The main issue for a use-case like NormCap are the trained models: they are optimized for images of _printed_ text and layouts, which is different from on-screen-text in many aspects. Unfortunately, I don't have the resources to train my own models.
Cuneiform was a long time competitor, but afaik development there is stalled.