Help Make the 19th Century Searchable
blog.archive.org
blog.archive.org
I think perhaps the developers took a wrong turn when they started trying to improve OCR with language models rather than font models. Humans can accurately transcribe a printed text without knowing the language. They do that with a mental model of font metrics and so on. In fact, a human transcribes more accurately (though more slowly) when they don't know the language because they don't erroneously "autocorrect".
I wonder if Google has in-house OCR that works much better than commercially available OCR. Google has OCR-ed at least 25 million books. You can't download the complete texts, but you can see snippets. Perhaps someone would like to publish a paper assessing the quality of Google's OCR and comparing it with commercial software. (Probably someone has done that already; I'm just bad at finding papers.)
THE NORTH STAR.
ee ee
ROCHESTER, DECEMBER 3, 1817.
— — ——--
‘THE COLORED CONVENTION,
We give Mr. Nell’s report of the doings of
this Convention, as the best we have scen. |
The crowded state of our columns prevent:
our publishing in the present number, any of
the able and interesting reports which en-.
gaged the attention of that body. We shall
attead to them in our next.
For the confidence reposed in me,You can feed it an image that contains text, with the only constrained being that the text should not be rotated. Depending how much text there is it can digitize a 12MP image in about .5 to 2 seconds on an iPhone Xs.
I compared it to ABBYYs OCR SDK and i have to say Apple's OCR outperforms it nearly every time both in speed as well as in recognition quality. I am still fascinated by it and would really like to know how this is working under the hood.
We found the document (the observatory yearbook written in 1859) via Google Books. As I don't speak Italian, I initially typed out passages into Google Translate in order to find the information I needed. Google Books has a view text option, but the format of the page and font often made it garbled (when pushed through translate in any case). Decent OCR likely would have made my life a lot easier.
There is a recent trend in space weather research to study extreme geomagnetic storms that happened in the 19th and early 20th century, and is aided partly by all of the scanned documents from that era available on the likes of Google Books, The Internet Archive and HaithiTrust. Better OCR would be a great help.
Although even having access to all of the documents is already incredible!
[1] https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/201...
https://arxiv.org/abs/2005.01583
https://news-navigator.labs.loc.gov/
What they have done already is great:
https://chroniclingamerica.loc.gov/newspapers/
I’ve used it quite a bit and they have some of the best OCR results I’ve ever encountered when it comes to scanned newspapers. But it’s just the tip if the iceberg. Historical newspapers seems to be one of the largest corpuses of material yet to be digitized.
Re their statement: "What we do not have is a good way to integrate work on these projects with the Internet Archive’s processing flow. So we need help and ideas there as well."
... maybe there are still people from the Gutenberg project around, they used to be handling human-transcribed stuff at quite a volume. (Personally I'd rather type stuff in than 'repairing' bad OCR, that constant back-and-forth is just aggravating. I'd be glad to dig in on something personally interesting, like say a little-remembered expedition or research in an area of interest ... they could arrange stuff by subject that way and then put up wish-lists for volunteers. History is full of forgotten amazings.)
All text is recognized, and as far as I can tell, there are no errors.
There are a lot of difficulties with older text, but I would start with low-hanging fruit such as trying to use better tools. On top of that, you have the problem of making sense of the layout, fixing common typos, etc.
I'm not sure though whether the IA can get Google to OCR it for them for a budget they can afford. Likely they'd want OCR solutions that have a one-time cost, so volume based SAAS offerings won't work.
I guess that would take a tremendous amount of time, but it would be really cool to get an old newspaper article to look at, and I bet a lot of people would find interesting things to talk about.
You could have a "comment thread" for people who added the piece, kind of like adding a digital layer to history.
Besides, as you said, reading the old paper would be enough of a prize in itself.
Less happy about the latest one but the book transcription and map correction were good as I benefited from both those projects (used the tools).
[1] and good luck with Beowulf: http://www.bl.uk/manuscripts/Viewer.aspx?ref=cotton_ms_vitel...
Related to the overall culture tech, most of the scientific revolution occurred in NeoLatin which isn't taught, mostly remains untranslated and untranscribed. For instance: Descartes first book, which dealt with music theory and human emotion.
Not sure if I'm just not getting the full quality somehow or if the image quality just isn't there for OCR to ever work?
It also seems likely that the jpeg2000 versions are the original, so they may be slightly better, but haven't been able to open one of those yet.
I do like the idea and efforts from a technical point of view. Tinkering with OCR on unusual (or old) languages. But that's not the goal of this project as far as I'm concerned (it's a byproduct?) Archiving every single news entry for the sake of completion sounds more like obsession than purpose. We're creating so much information that it will be even harder to separate garbage from valuable information (you have to spend time reading the useless stuff before you can justify whether or not it's valuable to you).
Information overload IS a problem and by adding more information to an already saturated ecosystem I don't see the vision here but would like to understand :)
It almost seems like a hording problem but for the digital natives. People accumulate a lot of stuff but rarely can they actually appreciate what they possess as time & perception is a very limiting factor.
An article that would shed some light would be highly appreciated.
We're probably not that good at recognizing which bits will be of interest to future generations, so archiving everything (to a point ... but I believe that newspapers are well within reasonable) sounds like a good idea. Plus you never know what you discover when you make things available.
While this pales in comparison with the 6 billion humans added since, that's a significant change - particularly for most likely available recorded sources - at a time of monstrous evolutions to major world powers/empires, expansion into vast new areas of the world (namely North America) of essentially the British empire (while at the same time the East India Company ceased to exist by the end of the century for contrast), some abolition of slavery becoming a reality in places (1833 for the British), and what arguably kickstarted much of the mental frameworks for our entire lives: the first two industrial revolutions (for example: democratization of once-monastic school system while adopting the year-of-production type of mental model for its promotions).
1804 is the first locomotive. 1859 is The Origin of Species by Darwin. 1861 is Maxwell equations. 1869 is Mendeleev's period table. And so on and so forth[0]. Measurement devices also improve in reliability and efficiency, leading to many of the early recordings we can now look back at when it comes to the consequences of the explosion of human activity with regards to the environment.
It's quite a fantastic century to keep a trace of, frankly.
[0] https://en.wikipedia.org/wiki/19th_century#Science_and_techn...