Know of any classic OCR tools that can reliably extract tabular data from scrappy PDFs? I've been hunting for a dependable option for that for years.
It's a very hard space in the long tail, like tables that span pages or tables with complex internal structures. I went into it thinking "eh how hard can tables be?". Very hard. Thankfully it's a pretty active research area.