Fine-tuning does help, though.
These are some of the nuances we had to work with during VLM fine-tuning with structured JSON.
What you can do is for instance to ask one model for a transcription, and ask a second model to compare the transcription to the image and correct any errors it finds. You actually have a lot of budget to try things like these if the alternative is to fine-tune your own model.
Best 2 out of 3 should be far more reliable than any model on its own. You could even weight their responses for different types of results, like say model B is consistently better for serif fonts, maybe their confidence counts for 1.5 times as much as the confidence of models A and C.
It is an absolute miracle.
It is transmutating a picture into JSON.
I never thought this would be possible in my lifetime.
But that is different from what your interlocutor is discussing.
I used to work in Computer Vision and Image Processing. These days I utter this sentence on an almost daily basis. :-D