These kinds of errors have always existed and will always exist there is no perfect way to extract info from documents like this.
The open question is how much better they need to get before they can be deployed for situations like this that require a VERY high level of reliability.