Ask HN: Any PDF Benchmarks?
I am testing several PDF parsing libraries, and unfortunately, most of them have issues. For example, many struggle with non-English languages, while others can’t reliably handle tables. Having a standardized PDF benchmark would be incredibly helpful — at the very least, it would allow us to identify obvious shortcomings without needing to install and test each library individually.
Is anyone building such a benchmark?