Base64.ai – Extract text, data, photos and more from all types of docs
base64.ai
base64.ai
After looking at options and few tests, I figured I'd use https://github.com/jbarlow83/OCRmyPDF It converts the PDF to an image for Tesseract and then recreates the PDF with the text copy-able.
It won't identify the address part of a driver's license, but that wasn't necessary for this project.
I've been thinking of running OCR on video frames. I'd also like to do speech-to-text extraction for searching my archives later (have about 4TB of video to trawl through, and desire text-based search capabilities). It's an interesting space to explore, but everything's been moving to web-service at a cost-prohibitive model.
For speech to text.. if english, try mozilla's deepspeech? https://github.com/mozilla/DeepSpeech
Might be fun to try.
[0] https://stackoverflow.com/questions/27568254/how-to-extract-...
Thanks so much for the tip on DeepSpeech!
Free software (AGPL-3.0 License), fast, highly accurate and extremely simple to deploy (I have no affiliation with them).
Full disclosure im one of the founders.
At the time there were not may apps out there and we partnered with a 3rd party service who did the OCR off the app so our quality of conversion (at the time) was close to state of the art from a mobile once people got comfortable with this method (which of course not everyone did).
We made some decent money as a side project from it but I also started to appreciate the sheer complexity of OCR.
We spent a lot of time fine tuning pre-processing before hitting the OCR engine (e.g. orientation, shading) small changes here made huge impact to performance. We also built various prompts to guide the user on how to take the photo to help. Managing expectations was something we were very conscious off and it was tough.
The unexpected use (but rewarding) use case was when we found people who were blind started to use the app to help with their daily lives - only a few but it was making a real impact to them so we priortized a few features to this segment knowing we were drifting away from maximizing revenue but we were cool with this as it was not a primary income source.
In the end we all moved to other things, more apps / services came on the market, google lens became a thing so we decided to sunset the product and did our best to manage customers through this process.
A rewarding experience overall - lots of lessons were learn that I have used elsewhere in my life since and ticked off' Build an app that made thousands of $' of my bucket list (which yea I should probably review!).
I'm assuming they only trained on some specific documents (passport of country X, etc) and all others don't work.
If someone processes the same document all the time, then my invoice2data project may work better and is open source. It's based on Regx, rather than machine learning: https://github.com/invoice-x/invoice2data
What was the process resulting in this name?
Edit: I am not being critical, I am really asking.
> Base64.ai SOC 2 compliancecertifies our bank-level security standards. Our API does not store your data to prevent possible data breaches. All API traffic must be authenticated and encrypted over HTTPS.
Sounds... Good enough? I mean, for what it is, it sounds like it's at least trying.
https://www.acuant.com/idscan-data-capture-software/
We also offer products that they don't provide. Our AI is capable of analyzing sound data (speech to text). It is extensible to add your custom forms and document types. We provide a cloud API and RPA components for UiPath, Bardeen and other RPA providers. We built Base64.ai so that you won't need a new vendor for new document types and platforms.
Happy to meet over Zoom if you want to learn more https://base64.ai/meeting
Manual data entry for an _entire page_ of text is about 15c, or 10c at volume.
We are a pure AI company, i.e. there is no human-in-the-loop. We are and strive to be more accurate than manual labor, and our processing time is 1 second rather than minutes-to-hours. Also our AI is naturally unbiased and does not discriminate.
In that case, shouldn't it a fraction of a penny rather than a whole dollar? Automation is supposed mean lower costs.
Sending such sensitive data "into the cloud" is no joke, for any company.
"First Name": "400 MHz ~2433,5MH:"
"Issuing authority": "0MHz~2833.5MH:"
may be the requirements about the documents the system can accurately recognize need to be explained in the app.