Anything that mentions tesseract is about 10 years out of date at this point.
Their training code and data is closed source. They are barely open weight and only inference is open source.
Using Surya gets you significantly better results and makes almost all the work detailed in the article largely unnecessary.
That doesn't hold for any of the GPU-based solutions, last time I checked.
VLLM hallucination is a blocker for my use case.
Otherwise I'd say just use your operating system's OCR API. Both Windows and MacOS have excellent APIs for this.