InvoiceNet: Neural network to extract information from invoice documents
github.com
github.com
Coincidentally I'm just about to begin a project intending to use the Form Recognizer service in Azure:
https://azure.microsoft.com/en-us/services/cognitive-service...
I will definitely do a side-by-side comparison with InvoiceNet.
Xero also has a related service: https://www.xero.com/au/features-and-tools/accounting-softwa...
We will open our API next week at last.
That's why we're looking at doing this in-house (via Azure or something ala this) even though we really don't want to.
Invoice recognition is a tricky subject, the companies that specialize in this field have spent a large amount of time and money on the problem, it would be great to see some kind of benchmark vs the commercial services.
Addressed in the disclaimer section. :)
For each form type that you intend to extract data from you will need a substantial training database. The good news is as you use it you build up more data but for legal reasons you may not be able to use that data to train on.
So that's why it really needs a dataset. One way to get one is to generate it based on a real dataset. I think that stands a much higher chance of happening than that some company will ship their - highly confidential - invoice stack to an unknown entity to make it world readable. That would likely cause that company serious problems and their legal department would never sign off on it.
For example, what if I embed the invoice data in a JSON file in a PDF? Could that make it easier for the user?
I really don't know much about PDF, but from what little I just read after checking it is possible to do that.
Here is a python lib that does it: https://pypi.org/project/factur-x/
Odoo, an open source invoicing software, produces factur-x invoices systematically (whatever the country): very convenient as it's parsed automatically.
Another benefit is that it's really easy to recognize an invoice at a glance, and you know exactly where all the info is.
[1] https://en.wikipedia.org/wiki/Universal_Business_Language
If you have access to a large number of vendors (invoice templates), these templates are going to be really valuable for creating better models. If you provide samples as a free/open dataset, you could make all invoice models that use your dataset, even if developed by other companies, better on your distribution.
js https://www.npmjs.com/package/pdf2json
py https://py-pdf-parser.readthedocs.io/en/latest/ or https://pypi.org/project/pdfminer/
Unless the training data is an extremely diverse set of invoices, maybe randomly generated?
The problem with invoices is that some fields have extreme variability - the address, the company names and the product descriptions. So a synthetic invoice generation approach might not work when you want to process in a new industry or language.
FWIW, one can test the AI of Odoo here: https://www.odoo.com/page/invoice-automation