Around couple of years ago I am working on a home project and utilised Tesseract and Laptonica for OCR. Storage and search HDFS, HBase and SolrCloud on extracted text. You can find the details here on my website. I was very impressed with conversion of hand written pdf docs with 90% readable accuracy. I have named it as Content Data Store(CDS) http://ammozon.co.in/headtohead/?p=153 . Source code is open and you may find steps on installation and how to run here.
http://ammozon.co.in/headtohead/?p=129
http://ammozon.co.in/headtohead/?p=126
A short demo
http://ammozon.co.in/gif/ocr.gif
I didnot get time to enhance it further but planning to containerize the whole application. See if you find it useful in its current form.