In real world usage, many tables are badly misaligned. Headers are off. Lines are missing between rows. Some columns and rows are separated by colors. Cells are merged. Some are imported from Excel. There are dotted sub sections, tables inside cells etc. Claude (and now Gemini) can parse complex tables and convert that to meaningful data. Your solution will likely fail, because rules are fuzzy in the same way written language is fuzzy.
Recently someone posted this on HN, it's a good read: https://lukaspetersson.com/blog/2025/bitter-vertical/
> You don't use multimodal models to extract a wall of text from an image. They hallucinate constantly the second you get past perfect 100% high-fidelity images.
No, not like that, but often as nested Json or Xml. For financial documents, our accuracy was above 99%. There are many ways to do error checking to figure out which ones are likely to have errors.
> This is using exactly the wrong tools at every stage of the OCR pipeline, and the cost is astronomical as a result.
One should refrain making statements about cost without knowing how and where it'll be used. When processing millions of PDFs, it could be a problem. When processing 1000, one might prefer Gemini/other over spending engineering time. There are many apps where processing a single doc is say $10 in revenue. You don't care about OCR costs.
> I've build a system that read 500k pages _per day_ using the above completely locally on a machine that cost $20k.
The author presented techniques which worked for them. It may not work for you, because there's no one-size-fits-all for these kinds of problems.