I'm curious if a multimodal model would be better at the OCR step than tesseract? Probably would increase the cost but I wonder if that would be offset by needing less post processing.
Florence-2 is < 2GB so it fits into RAM well, and it is MIT licensed!
On a T4 in Colab, you can run inference in < 1s per image.
[edit] But it is not applicable to OCR specialised models like Florence-2