As in: you tell it that these and these parts should be masked such and such, and then it does that?
Edit: Thank you!
However Tesseract is quite behind still as you note, even with v5.
I would say PaddleOCR is good in general for tables - it's much better (in terms of recall rate) at recognising numerical digits / symbols than Tesseract although I notice it often misrecognises "l" in "Lullaby/ml/million" etc as "1" sometimes.
The cloud providers have better table extraction iff you can guarantee the same format each time for the document.
Here's just one example: https://www.totalflood.com/samples/residential.pdf (I struggle getting accurate data out of the Sales Comp section - basically all approaches mix up the properties.
I would suggest your best bet is waiting 2 years for the next version of LLAVA to come out which may have capabilities to interpret very accurately on device. The progress with LLAVA has been fast recently but for now it's still a bit too inaccurate.
Does this require a slight offset and/or rotation to the image, or just literal rerun, with seed seed/whatever giving a different result?
And these haven't even been trained to defeat captchas/logic problem captchas yet, if it was fine tuned on the general pattern of them I imagine any form of captcha is bust.
I don't see how an older smartphone could meaningfully outcompute a spamming infra.
Of course eventually this will be defeated too, but for now it seems to work pretty well.