Show HN: BetterOCR combines and corrects multiple OCR engines with an LLM
github.com
github.com
Any plans to add other OCR engines?
- https://github.com/mindee/doctr
- https://github.com/open-mmlab/mmocr
- https://github.com/PaddlePaddle/PaddleOCR (honestly I don't know Mandarin so I'm a bit stuck)
- https://github.com/clovaai/donut -- While it's primarily an "OCR-free document understanding transformer," I think it's worth experimenting with. Think I can sort this out by letting the LLM reason through it multiple times (although this will impact performance)
- yesterday got a suggestion to consider https://github.com/kakaobrain/pororo -- don't think development is still active but the results are pretty great on Korean text
Looks great! I’ll give it a go.
If you can’t highlight the text, it won’t work.
PDF -> Markdown looks like a pretty great use case
Just added box detection support -- maybe I'll start from here https://github.com/junhoyeo/BetterOCR#-box-detection
It does what you are looking for