The other thing I'm interested in, and wonder if you have thoughts on: When you pass, your personal library is inherited. Will people know what to do with your digital library? Will it be useful? Will people want it?
The other thing I'm interested in, and wonder if you have thoughts on: When you pass, your personal library is inherited. Will people know what to do with your digital library? Will it be useful? Will people want it?
The tools for processing PDFs into searchable text have a lot of warts. For a while IBM was offering a free Watson service to do this (now its part of Watson Discovery) which has some warts. I did manage a set of perl scripts that would post process the statements that I downloaded from the bank into CSV files, but I would still like to pull tabular data out of PDF book scans to make the data they provide more useful.
I have a simple frontend based on the perl Mojolicious module which Blekko had developed as part of another project but my indexing tools are still quite primitive. Simple bi-gram and tri-grams, and a growing synonym index. I don't give it enough queries to use my own traffic for ranking feedback. So basically everything is nearly equal rank. Basically I am about to the AltaVista level of search capability :-).
The vision is it just runs as a server and anyone on the same network can access it like a web service an pull up documents (and in the future media) of interest.