I currently use https://github.com/firecrawl/anydoc in my PDF pipelines.
In most cases it performs well. I'll test your lib to compare.
112 karma · joined August 24, 2026
After running Google Ads for a couple of years, our current exclusion list has over 4000 networks just in the US.
Local: https://github.com/datalab-to/chandra Hosted: https://www.datalab.to
Another decent option is GLM OCR. It's slightly less accurate but faster and cheaper.
Local: https://github.com/zai-org/GLM-OCR Hosted: https://docs.z.ai/guides/vlm/glm-ocr
Other models such as PaddleOCR, dots.ocr and DeepSeek OCR performed significantly worse.