Tika's PDF text extraction is fine if you're just trying to get searchable text. Which is what it's made for: Slurping doucments into Lucene. Fulltext search typically isn't terribly sensitive to getting the order of words right, and is even less sensitive to getting the formatting right.
If you're trying to get something fit for consumption by human (including via a screen reader) or an NLP pipeline, though, all the problems discussed in that FilingDB article still apply.