There is so much potential to these technologies, but even autogenerated documents have so many embedded semantics that it is hard to train these tools for a wide variety of document formats.
(After all, a human can do it, with some basic training. Ergo we should aim at computer being able to do it with equally simple training too. And in practice this doesn't seem like an insurmountable goal.)
The French standard is called Factur-X. The German one ZUGFeRD.
More details here:
https://www.pdflib.com/pdf-knowledge-base/zugferd-and-factur...