I find it challenging to accept something that talks about "OCR" but then I upload a PDF with text in images, and when I query the document after upload, I get a message that says "I can't interpret images"..
Then are you actually doing OCR, or are you just extracting embedded text?