Probabilistic Scraping of Plain Text Tables (2013)
edinburghhacklab.com
edinburghhacklab.com
Would you mind sharing how many rows was it? I know it's probably possible to reverse-engineer it from other information, but I think it would be illustrative.
Also, this fragment is interesting:
>Whilst integer programming is NP-hard to solve in general, these problem instances are not pathological instances
Could this kind of tabular data have turned out to be pathological? Would it mean that constraints cannot be met and we had to search the whole space to ascertain that? I imagine these general solvers don't do specific heuristics when searching.
> these problem instances are not pathological instances
With MIP, the solver churns when you get a lot of information suggesting that one path is likely to be the optimal but actually it's another path but you have been mislead by red herrings. One cause would be a subset of weak classifiers are configured incorrectly so they are actively misleading. Another cause might be that the table structure is very ambiguous and requires a lot of global deductive reasoning to figure it out.
However, given digikey is not maliciously trying to create difficult to read tables, I don't think these cases really come up. Misleading weak classifiers would be a bug, though its sometimes hard to spot if enough of the system still reaches the correct decision. There is probably some math to assign utility to certain classifiers though I never looked at it beyond case by case debugging of decision samples.
Also worth noting this never went into full production due to the surrounding business being broken for completely non-technical reasons. But I do think the MIP + MLE is a useful technique for a few different forms of problems where you want to integrate hard ontological constraints over a fuzzy reasoning system.
Or, was it due data context, which I assume is more plausible. I ask because I maintain an old, slightly large, and growing project which contains about 12 different regexes and is used on messy unstructured data. I'm in the process of rewriting it into a more general framework using NER + RNNs or HMMs, but this seems like a very interesting approach.