I'll show myself out
2,328 karma · joined September 29, 2009
I'll show myself out
We're super flattered over all the attention and we're listening for any feedback.
It's designed to work primarily with old fire insurance atlases (e.g. Sanborns) and is a bunch of hand-tuned heuristics. But it powers Building Inspector (http://buildinginspector.nypl.org), which is where all the data is validated consensus crowdsourcing, providing ground truth for the data coming out.
Unfortunately (or perhaps fortunately if you're trying to get a Computer Science PhD) I haven't come across anyone applying deep learning to map vectorization either. Frankly, it could open up a whole new field with respect to historical mapping (among other things).
Would love to talk more about it and see how deep learning could be applied here. We always structured the outputs of Building Inspector so they'd be useful training sets for unsupervised or reinforcement learning so it could hopefully apply here as well.
A lot of this comes down to how much time you're willing to put into it to get it up and running and if you're willing to put in any effort to additionally train/refine a model.
(ps all the elements I'm about to mention here have been mentioned throughout this thread. I'm just providing a bit more context and bringing what I know together).
_________________
Gotta have something now: Tesseract
Not saying it's "bad" but it ain't as great as you want. Bindings for every language, lots of configs, documented well enough that it's 80% of the suggestions here. For what it is, it's a damn miracle. Most languages and alphabets and lots of kinds of fonts likely start giving you results somewhere between 80% and the Uncanny Valley immediately. You also get a pretty good Layout Analysis engine [1] for if you're working with complete pages of text. The existing models it ships with are robust, but if you want to get better outputs, retraining it is a real pain (you have to segment on a character-by-character basis). You're better off trying to apply a general preprocessing to clean the image or...
__________________
Gotta have a proof of concept this afternoon: Send just text to Tesseract
By now you've realized that asking an OCR engine to OCR a tree (or a picture of a tree alongside some text) somehow always comes back looking like a cat just napped on your keyboard. As alluded to in a few other places here, Just Give It Text! Tesseract (and most traditional OCR) was developed in a text document centric world where you could safely assume it's just a page of words, and sometimes you might have to deal with columns (that's where the Layout Analysis comes in). Probably not your real world.
Depending on your type of documents you might be able to develop a few hueristics for identifying text regions in the picture, then sending only those sections over to Tesseract. This'll dramaticlly help Tesseract out, but it can increase some of the complexity you'll have to juggle. You might be able to come up with some hueristics specific to your documents (which can be very good, especially if it lets you infer more information about those regions that you might want later).
You can also use something like Stroke-Width Transform (all the links I would use were graciously linked in this earlier comment [2]) which was discovered by Microsoft trying to spot text in the wild for their Street View efforts. ccv has a very nice SWT implementation [3], and their http server for the whole library [4] that with a bit of makefile finagleing can have a very nice SWT preprocessor -> Tesseract api over HTTP up in an hour or so.
Also it looks like the OpenOCR[5] project is now using SWT -> Tesseract with a nice HTTP API written in Go, conveniently packaged up for Docker and very well documented[6].
_____________________
Take me down the rabbit hole, but maybe get results in 2 days:
The roots of tesseract were planted by HP in 1985, so there had to be a better way at some point. OCRopus[7] was supposed to be the great state-of-the-art hope that would save us all. The approach was incredible and the results being published were great. But the documentation came in the form of a mind map[8]. Recently the project was picked up again by the original developer and rechristened OCRopy[9] and has garnered a pretty active and growing community in the past few months.
You'll have to do a bit more work here than just tesseract, but the LSTM neural network[10] approach completely blows away Tesseract's results with just a little training.
To get started, Dan Vanderkam's tutorial is excellent to start working with the out-of-the-box model[11] immediately.
But results get INCREDIBLE when you take some time to train your own model[12]. I provided Dan with the source images for the project[13], from The New York Public Library's historical collections, and the source text was in a font style the default model had never encountered before. But in an hour or two of transcribing the training data (training data in OCRopus is awesome because you just feed in a line at a time; no need to align and segment each letter like in Tesseract), he was getting an under 1% error rate!
Layout analysis isn't OCRopy's strongest suit, so you might get even better results if you pre-segment with something like SWT, but again, not completely necessary.
____________
That's it. Way too much for a Newsy Combinator comment, but a pretty decent tour of the big stuff in open ocr world. There's some magic state of the art going on inside Google, Microsoft, and Abbyy, but hopefully you'll be able to teach your computer to appreciate whatever it is you want it to read through these now.
[0]: http://nypl.org/labs
[1]: https://en.wikipedia.org/wiki/Document_layout_analysis
[2]: https://news.ycombinator.com/item?id=9776034
[3]: http://libccv.org/doc/doc-swt/ and http://libccv.org/lib/ccv-swt/
[4]: http://libccv.org/doc/doc-http/ and http://libccv.org/lib/ccv-serve/
[5]: https://github.com/tleyden/open-ocr
[6]: https://github.com/tleyden/open-ocr/wiki/Stroke-Width-Transf...
[7]: https://en.wikipedia.org/wiki/OCRopus
[8]: https://www.mindmeister.com/192257150/ocropus-overview-publi...
[9]: https://github.com/tmbdev/ocropy | all day all week
[10]: https://en.wikipedia.org/wiki/Long_short_term_memory
[11]: http://www.danvk.org/2015/01/09/extracting-text-from-an-imag... | though you've got to download the model separately because it's too big for github
[12]: http://www.danvk.org/2015/01/11/training-an-ocropus-ocr-mode... [13]: https://www.oldnyc.org/
We've been so surprised at the excitement over the response and we've had to bump up our timeline to catch up the energy of everyone who's gotten in touch.
There's a few big things we're focused on right now: * Improving performance of CV recognition of historical buildings [https://github.com/nypl/map-vectorizer] from insurance atlases. * Parsing data from historical records. We love old documents that are basically databases in print form. Old playbills, city directories, tax records; lots of tabular data with locations, people, businesses - the names of history. Getting those names out and tied to locations are what become the searchable elements for places. * Rebuilding the underlying gazetteer. We worked with the team from Topomancy to prototype the underlying system a few years ago [https://github.com/topomancy/gazetteer]
There's a lot more, but that's where were starting. If you're interested, drop us a line at spacetime@nypl.org or sign up for our (comically nascent) forum at http://talk.spacetime.nypl.org where we're trying to coordinate most of the effort.
For ebooks with images [Gutenberg]: https://github.com/mgiraldo/tenprintcover-p5
As for the dates, most are guesses. If you plot the dates, you see they're overwhelmingly years that end in 0.
And while not mapping, John Resig's Ukiyo-e [2] project is one of the coolest projects I've seen in the digital cultural heritage space in some time. He's been applying image recognition to these incredible Japanese woodblock prints from museums, galleries, dealers, universities and libraries all around the world. Because these things are prints, there could be hundreds of prints from the same block master all around the world, but because the expertise in the field is so divergent, the cataloging practices are really inconsistent. Different institutions might call artists by totally different names (or think a print is by totally different artists). So he built a search-by-image search engine of hundreds of thousands of Ukiyo-e that finds and reunifies prints totally independently of their metadata. Which is cool when you find 5-10 of the same print in places around the world. But it's cooler when it matches 2 prints with totally different artists and publishers and dates because at some point after the first print was made, someone bought the block master, cut out the face and replaced it with another, then did the same for the signature [3][4].
Actually, The Vectorizer owes a big debt of gratitude to John and his brother Mike. Mike Resig is a geographer and was the first to show us a process for how this kind of automated identification is possible.
[1]: http://www.mapbox.com/blog/user-friendly-guided-feature-extr... [2]: http://ukiyo-e.org [3]: http://ukiyo-e.org/image/met/DP134583 [4]: http://ukiyo-e.org/image/mfa/sc214530
We, however are far more lenient (mostly because we can't afford to build a time machine to map the past ourselves).
On the other hand, Mike Migurski's Green Means Go [1] project is fantastic for figuring out where batch imports into OSM will be greeted with confetti and parades for filling out parts of the US without enough coverage to warrant anti-import protectionism.
These are nice to have for researchers, but the real purpose of collecting this is, just as you note, to unlock the hidden historical geospatial data in textual materials. Once we've got all those names of places, their addresses, their lat/lon coordinates, and their timeframes of existence, we can start to search through texts to find linkages. Old city directories (they're basically books of ghosts) start to show you who lived and worked where [1] (and in the process starts to get you more names you can associate with these places), address matches in historical newspapers start to show you what happened in these places, and the maps start to become this geospatial backbone to traverse across tons of different datasets.
The Vectorizer is so freaking cool for so many reasons, but mostly because it's going to let us actually get through these insurance atlases to collect this data before we all die (one of our favorites is the 1854 William Perris Atlas [2][3] but it took nearly 3 years to actually get through the 64,000+ buildings in Manhattan south of 42nd st) so we can start doing this kind of querying with it. The real geniuses behind all this, our Geospatial Librarian Matt Knutzen and the team at Topomancy, have been working on an experimental gazetteer [4] so that we'll finally have this as a public web service for people to hack on all these places as we collect and conflate them. Give us a few months...
In the meantime, sign up for the Open Historical Maps project listserv [5] that some of the OSM crew is working on (including the geniuses at Topomancy).
Also, this came out of a historical geospatial hack day [5] we threw a few months back, which you should check out if you want to play around with some of our data sources for this kind of work or for building something else out of historical NYC's geospatial footprint.
[1]: http://andrewxhill.github.io/cartodb-examples/scroll-story/b... [2]: http://maps.nypl.org/warper/layers/861 Tileserver, please forgive me for linking to you [3]: http://aaronland.info/nypl-perris/ YEAH SHAPEFILES! [4]: http://vimeopro.com/openstreetmapus/state-of-the-map-us-2013... Schuyler Earle's presentation on their version of historical gazetteer they're building for the Library of Congress at State of The Map US 2013 [5]: http://www.nypl.org/blog/2013/07/12/maphack-hacking-nycs-pas...
I'm also looking forward to people competitively racing these things with the personal trip data a part of the platform.
Aaron knew about copyright's implications - hell right around the time of RSS, he helped launch Creative Commons. And he didn't put a license on his site. So maybe by not stating a license he was making a statement about copyright. But now we'll never be able to know for sure.
It's been a while since the Interactive News Team profile a few years back; I think we'd all like to know how a major night goes down in the newsroom for your team.
We wonder what books scholars will write about 500 years from now. It won't be what's popular, it'll be whats pirated.
One of DirecTV's core claims has been that online streaming of TV content undermines the value proposition of for-pay television [1]. I get that having your most vocal, most internetty users channel their anger toward your opposition is a powerful negotiation tactic, but this seems like it has a strong potential to backfire against Viacom.
[1] http://mediadecoder.blogs.nytimes.com/2012/07/10/directv-via...
Of course, this is a concept hackers have understood for decades. It's why Jason Scott collects text files, scans old arcade manuals, and why Archive Team rips the entirety of Yahoo! Video, Geocities, and poetry communities just before they're about to be shut down. And they then ensure sure they're available to all who ask. For free. Forever. On Bittorrent.
Unfortunately, the old stuff isn't as easy to copy. It's locked up in big places like my office, The New York Public Library. But cultural memory organizations, even if they don't realize it, are WHERE ITS AT when it comes to transforming all that documentation of human endeavor and into some of the most powerful, useful, and amazing materials going forward.
While we used to treat our super rare special collections as cherished things never to see the light of day except by a couple of curators, interns and a few handpicked researchers and only then in climate controlled bunkers. On rare occasions, the public sees them only as part of an exhibition, on display once every few years for a few hours a day. While this is great for the insuring the longevity of the object, it's not so great for exposing the public to tangible history. Now we're opening up the rarest of materials, which are often the freest in terms of restrictions (due to their age), making them public, and making them richer with use. Suddenly using the rare stuff makes it better with every use. Way better.
These kinds of places are among the best beachheads in the fight for Free Culture. Some of the smartest thinkers about the copyright fight are now leading the Harvard Libraries from the inside. John Palfry's the head of the Harvard Law Library (for a few more days), David Weinberger is leading the team that's writing the code behind the Digital Public Library of America, and Jonathan (JZ) Zittrain, who cofounded the Berkman Center for Internet and Society, is about to become the Law Library's Vice-Dean.
Basically what I'm saying is that there are vast swaths of human knowledge and endeavor at crazy risk of being left behind. This isn't bit rot. This is real rot. Or sometimes we "digitize" something by just taking a picture of it. But then it's just a picture, not a page from an atlas with grounding in a a real time and place, and it gets lost in a photo gallery and a metadata system. And it's just sitting there, waiting for someone to come along and breathe life back into it, to make it something more powerful than it ever was on paper. To be fully indexed, crosslinked, talking to all the other information available, related, and relevant. That's eminently possible, it's just not being done at sufficient scale.
We need hackers in every position in these organizations. Here at the New York Public Library, we're working on that. We have one curator who hacks on Node and Backbone because it means he can better disseminate and preserve musical theatre (also because it's really fun). We have one curator who's digitizing copyright-free maps of New York City and is using it to basically create historical Google Maps all the way back to when Europeans first showed up on North American shores - or as someone just called it "Building guidebooks for time travelers" [1] (I'm particularly proud of that one). And now I'm going to move into the realm of shameless self-promotion and a plea. My team looks for cool datasets throughout the library and builds projects on top to transform them into structured data by any means necessary (often it's often crowdsourcing) then builds APIs so anyone can get to that data. We also build animated gif creators [2].
Basically what I'm saying is that those rare books, those archives, those old photographs need hackers to keep these materials around. To keep them relevant. To inspire countless others to create great works. And here's the point. We need so many more people helping us.
We need you. Not just at NYPL, but everywhere across this industry. Because how cool is it to hack on the sum total of human knowledge and make it available to all who ask? If this piques your interest, consider joining our team[3]. (Or work on this issues on your own and tell us about them.)
[1]http://betabeat.com/2012/06/doctor-who-companion-application... [2]http://stereo.nypl.org [3] https://jobs-nypl.icims.com/jobs/6882/job | We're not always hiring, but we've got an opening now if you want to join us
The notion that everything is on the record is of particular note. You, run into a journalist at a party and mention drunkenly your product launch strategy. Consider it out of the bag because it's fair game to report now. Expect a call the next morning.
Moreover that extends to anyone with familiarity with the situation. Your dirty laundry can be aired in just such a way. Your employees offhanded comments are fair game for stories.
I will note that the one thing you can and should do is contact a reporter's editor if they've run a story about your company without calling for comment. Be totally sober about it and make sure you've run whatever you plan to say by someone with emotional distance from the situation. It's easy for it to come off as "WAH WAH WAH why didn't I get called to talk about this bad thing" and the best way for you to get your side across is to be as detached as possible.
Yahoo News is still the top news site. Yahoo business/finance is still among the strongest financial properties. The good news is that Yahoo's not going anywhere. The bad news is they're not going anywhere.
Anyone know of a good tool to call in and rebroadcast the call so everyone can listen?
Under the EB-5 program for investors, by investing $500k in a targeted high unemployment area in a capital intensive project that creates jobs, an investor can get a green card.
I had no idea this program existed until this NYTimes program on ski resort Jay Peak, which raised $250m through the program. [http://travel.nytimes.com/2011/12/11/travel/with-renovations...]
Now if we'd only apply to entrepreneurs what we're already doing for investors...
Louie's just doing for comedy what have been doing for years, recording the concert and selling it at the end of the night or online for more than they'd have made otherwise through traditional distribution deals.
The takeaway is, unless these letters are hand delivered, I doubt theyll reach their intended recipients in time.
I've never heard a song that better sums up the feeling of building something incredible.