In-Browser OCR
ian-nai.github.io
ian-nai.github.io
I just tried it with the not-great-quality image at https://www.geograph.org.uk/photo/1587043 . The result for the first four lines was:
“This building was erected around 1874 to provide a location for a seismometer of the British Association. A seismometer is an instrument which is designed to record earthquarkes and the one located in this building was only one of a series of such instruments located in the vicinity of Comnie to investigate the earthquakes which had been, and continue to be, prevalent in the area.”
The only mistake seems to be “Comnie” instead of “Comrie.” (The misspelling “earthquarkes” appears in the original.)
In contrast, the In-Browser OCR gives the following for the same four lines:
“m .me W m m: swm mm m mm .. lucauan m a smmmm. m m. anus» Lssm'm'm» A swmmm ‘5 3,. WWW mm. vs damned .a mom “mum: m m: an» army! m a... “mm was Only an. a: .1 gem a; 5m m<llumems mm m we mm 0. Camus w WNW: we earmquakes Wm»... mm W. m coulmu: m be Wax/Nam m m M”
> EARTHQUAKE HOUSE
> ‘his building was erected around 1874 to provide a location for a seismometer of the British Association. A seismometer is an instrument which is designed to record earthquarkes and the ‘one located in this building was only one of a series of such instruments located in the vicinity ‘earthquakes which had been, and continue to be, prevalent in the area.
> of Comrie to investigate
Clearly, Tesseract 4.0 has problems following the baseline of the text. But otherwise, it is much better than the output from the website and even got the title correct. Which makes me think they use an older version (?)
My commandline:
tesseract 1587043_dcd093c4.jpg output -l engThe 4.0 version added new neural network system based on LSTMs, with major accuracy gains.
https://fossies.org/diffs/tesseract/4.0.0_vs_4.1.0/ChangeLog...
> ARTHQUAKE HOUSE 1
> Ths builing was erecied Ground 1874 to provide a location for a seismomater of the Bitish Aasociation. A sefémomatar is an nsirument which is designed 10 racord earthquarkes and the ne ocate n his buiding was only one of a seies of such instruments located in the vicinity of Coms tonvestgete the earthauakes which had been, and continue to be, prevalet n the area.
The 4.0.0 version looked better to me ...
Tesseracts expects PNG and outputs to text. I want the same PDF to be hidden overlaid with OCR text.
This free PDF online service does a decent job but offline would be better.
It comes with an API, so you can integrate it with your workflow.
I use PDF24.org which actually does everything great.
Anything offline?