The tools for processing PDFs into searchable text have a lot of warts. For a while IBM was offering a free Watson service to do this (now its part of Watson Discovery) which has some warts. I did manage a set of perl scripts that would post process the statements that I downloaded from the bank into CSV files, but I would still like to pull tabular data out of PDF book scans to make the data they provide more useful.
I have a simple frontend based on the perl Mojolicious module which Blekko had developed as part of another project but my indexing tools are still quite primitive. Simple bi-gram and tri-grams, and a growing synonym index. I don't give it enough queries to use my own traffic for ranking feedback. So basically everything is nearly equal rank. Basically I am about to the AltaVista level of search capability :-).
The vision is it just runs as a server and anyone on the same network can access it like a web service an pull up documents (and in the future media) of interest.