I would expect the deep learning approach to outperform traditional approaches in terms of accuracy, but it would be good to see accuracy vs. CPU / memory used, etc.
I would expect the deep learning approach to outperform traditional approaches in terms of accuracy, but it would be good to see accuracy vs. CPU / memory used, etc.
A better comparison would be against Tesseract or ABBYY FineReader.
EDIT: I wasn't aware that Tika now embeds Tesseract.[1] Still, it's a simple wrapper so the real comparison is against Tesseract.
https://blogs.dropbox.com/tech/2017/04/creating-a-modern-ocr...
I'm not sure what benefit they are getting from using machine learning for this other than "decide whether to try and process this file or not".
Tika + Tesseract seems to be able to do the heavy lifting they spent a lot of time talking about in that article.