I think the traditional approach to scanning and classifying without AI/ML is the way to go, for the next 5 years at very least.
I think the traditional approach to scanning and classifying without AI/ML is the way to go, for the next 5 years at very least.
For my use cases, this already beats all "traditional approaches" for at least a few month now. That's just inferring from when I first stumbled across it. No clue for how long it's been a thing.
Google Vision: 95.62% HW - 99.4% Typed
Amazon Texttract: 95.63% HW - 99.3% Typed
Azure: 95.9% HW - 98.1% Typed
Then if curious, TrOCR was the best FOSS solution at 79.5% HW and 97.4% Typed. (However it took roughly 200x longer than Tesseract which was 43% HW and 97.0% Typed)
I also tested paddleocr and keras ocr to round them all out.
At some point I really need to finish my project enough to write up some blog articles and post a bunch of code repos for others to use.
Leaking sensitive data of enterprise customers as training material for public recaptchas falls in that category.
Which it used OCR to produce digital text.
So one source of training data at least.
But this isn’t much help if you must classify images.
It uses Tesseract under the hood. Results tend to just be OK in my experience.
I'll note that when I put the tesseract output into chatgpt and prompted it saying it was ocr'd text and asking to clean it up, it worked very well.
My first time processing it, I used `ocrmypdf --redo-ocr` because it looked like there was some existing OCR. After processing, the OCR was crap because ocrmypdf didn't realize it was OCR but thought it was real text in the document that should be kept. This was fixable using `ocrmypdf --force-ocr`.
Before realizing this, I discovered that Tesseract 4 & 5 use a neural network-based recognition. I then came across this step-by-step guide on fine-tuning Tesseract for a specific document set: https://www.statworx.com/en/content-hub/blog/fine-tuning-tes...
I didn't end up following the fine-tuning process because at this point `ocrmypdf --force-ocr` worked excellently, but I thought the draw_box_file_data.py script from their example was particularly useful: https://gist.github.com/flaviut/d901be509425098645e4ae527a9e...
I did a presentation on the topic recently: https://clis-everywhere.k8s.best/16
I'll soon make the stack open source, but it shouldn't be hard to recreate given the inputs I've already provided.
This seems like exactly the kind of problem that will see rapid improvements as people point more LLMs at multimodal input.
Right now making predictions for ML capabilities on a five year timeframe seems foolhardy.