Natural Language Processing for the Working Programmer
nlpwp.org
nlpwp.org
Is there a missing link there? I don't know. So I took a few days and built an extremely simple NER (named entity recognition) engine and made it extremely easy for any programmer to begin using.
See http://blog.nerily.com/howto-train-your-own-modelset-for-you... . We'll see how NLP becomes more easily accessible with better tools to come over time.
NLTK is also free and open-source, with a liberal license (Apache), which I appreciate greatly.
Also, I don't understand what you mean by nltk being a library for building NLP solutions, rather than NLP-powered apps. Can you expand on that?
Re: NER, I found this gist (not my own) for a basic example of entity extraction with nltk: https://gist.github.com/322906/90dea659c04570757cccf0ce1e6d2... This looks pretty straight-forward to me. What NLP toolkit are you using for the NER service your Chrome extension calls, if not NLTK?
NER comes to mind, lots and lots and lots of toolkits for building up to NER, but very few that let submit English text and get back a list of people, places and things without having to virtually build my own NER system from scratch anyways.
Give me NER, Entity relationships (ER) and a couple kinds of sentiment analysis scoring (SA) (which can be jump started with decent NER) and I've pretty much exhausted 95% of what I'd ever want to do.
I really really really don't need yet another library to do sentence tokenization, term tokenization, tf counting and stemming. If I was building a free text indexer or bayesian filter or some such it might be useful, but I'm probably not, there are far better solutions to those domain than I'm likely to come up with, but there aren't for NER, ER and SA.
It's pretty straightforward to use their library to read a document and output an XML file containing NER data (and lots of other fun stuff).
For instance, from the sentence:
> World War II, or the Second World War (often abbreviated as WWII or WW2), was a global military conflict lasting from 1939 to 1945, which involved most of the world's nations, including all of the great powers, eventually forming two opposing military alliances, the Allies and Axis.
Stanford NLP NER will output the following entities:
"World War II" - MISC "Second World War" - MISC "1939 to 1945" - DATE - NORMALIZED 1939/1945 "Axis" - MISC
You can view the output of Stanford's CoreNLP library (NER + dependency grammar + coreference resolution + some other stuff) for the Wikipedia article on World War II in my github repo:
https://raw.github.com/ryantanner/thesis/master/data/ww2samp...
edit: I should add that the real fun (for me) came from combining NER with dependency grammars and coreference resolution. It makes it very easy to turn Stanford NLP's output into a knowledge graph combining a large number of documents.
This is a bit more human friendly.
If you need an example sentence: "Stanford University is located in California. It is a great university."
I also know that Microsoft Research has a demo online of their NLP tools: http://msrsplatdemo.cloudapp.net/ (Silverlight required) I don't think you can download the tools though, but they do offer to provide you with an API token to call their service from their cloud.
Potential conflict of interest: I wrote parts of the CoreNLP visualiser.
Edit: Added example sentence.
There are other text processing APIs on there as well. As for libraries, I primarily come from the JVM camp for NLP, but I would recommend the following libraries:
http://nlp.stanford.edu/software/index.shtml (Comprehensive) http://code.google.com/p/clearnlp/ (Fairly simple)
My favorite is cleartk (http://code.google.com/p/cleartk/ ) mainly due to the fact it's a consistent interface, but UIMA itself can be a difficult toolchain to pick up, and I could understand most of these being overkill for many simple applications people may have in mind.
but haven't had the time to look at it with any detail.
The big problem I think with NLP in general is typically to do anything, you need a pipeline. (Sentence segmentation, tokenization, part of speech tagging) usually at a bare minimum. Then from there you can do named entity recognition or other tasks that produce actual usable results.
The NLP systems leading in competitions such as CONLL conference tend to be publicly available, so you can get a "general purpose" system there; but the current way usually is to train specific model for a specific purpose - since if you don't have a predetermined purpose, you can't really tell which of items should be tagged as places (instead of things); you can't tell which things should be tagged as 'things' and in what way they should be classified deeper - the list of classes tends to be application-specific.
So, the book is basically frozen. We hope to have more time in the future to continue the writing...
That was the first book in NLP (and the only for now) that I read. I've been interested both in NLP and Haskell. In that respect it fitted, thanks!
A few points to criticize. For the frequency list one should use multisets, not dictionaries. There are a few multiset packages at Hackage. Suffix arrays are badly explained. Monads - very badly. With tagging there was an impression that it could be explained simpler.
Many things are announced but not touched. The book is not a book in fact, it's more like an article. Perhaps reconsider it in that way? But oke, hopefully you will find time to continue it as a book.
Perhaps meanwhile you can recommend some other book to continue reading on NLP?
Foundations of Statistical Natural Language Processing by Manning and Schütze
I have to say comments like yours are not really encouraging to continue writing ;).
Myself being in industry, I know how hard, near to impossible it is to find time for anything extra than work and family. And a decent book requires approximately the same amount of effort as finishing PhD. Perhaps that was my frustration coming out of the projects I had to abandon. :(
Thanks for refs!
As soon as you get a corpus of any reasonable size (and you'll have to use large corpora for any meaningful, non-toy results), the various Haskell String-like classes and laziness-control options are mandatory, but tricky/ugly when starting to use them.
You might try to understand monoids first, because you already have familiarity with many applications of monoids. The realization "oh, this is just two functions" at the heart of monoids is also what's at the heart of monads, but the applications are different.
Desugaring helps a lot too. I learned by avoiding do-notation, but you can learn do-notation at the same time if you try desugaring as you go, so you can make explicit what's going on under the covers.
It's like math, you have to keep playing with it until you grok it. I find it's good to try out different expressions in ghci, use :t a lot to see what types are coming back, to build intuition.
If you put a few hours into it for a few days in a row, you can probably get to this epiphany in one weekend. The trick is building up enough examples that your brain can generalize it. Nobody's going to learn it by staring at the abstract form and thinking hard--if we did work that way, there would be a lot more use of comonads. That's why it's important to re-type examples. You're not going to be able to write the examples yourself until you understand them, but working through them gives you something to build on, and builds healthy expectations (I'm going to need return here, because the naked value isn't in the monad, etc.)
The epiphany is worth it--but don't count on a monad tutorial to help you much, they're mainly a side-effect of other people having the epiphany.
Scala (or Java) is another great NLP language. It's got decent libraries (openNLP, mallet, mahout), hadoop, and Scala is almost as nice as Haskell.
Haskell's mechanisms for defining parsers, lexers, and other pattern match tools is so good it probably passes over the line from "pretty" to "objectively better".
A lot of people who need to lex and parse data and then act on it turn to Haskell. It has some really remarkable and efficient libraries. And even for "common" target languages it's reasonable to write extremely fast parsers. With tuning, projects like Aeson are among some of the fastest JSON parsers and writers out there (only a few projects exceed its speed and resource efficiency in ANY runtime).
The very same patterns that define "packrat-like" parsers (which share a strong relationship to the monadic and "arrow-adic" parsers) can be extended to define things like DFAs and semantic pattern matching. And languages with support for rich, somewhat lazy pattern matching like Haskell and Prolog wipe the floor with eager languages without (e.g., C), which is ideal for semantic analysis.
While not an "authority" in the subject, I've spent a lot of time working with some very skilled folks in the field of NLP, Linguistics. Most tools they used (in our case licensed from X/PARC) had C underpinnings for performance, but ultimately consumed specifications that were very much like Prolog or Haskell in character. Talking to some of the linguists who wrote those tools suggested that had GHC existed (or Allegro or a fast prolog been cheaper) then they would have been much easier to write in those languages.
http://en.wikipedia.org/wiki/Frederick_Jelinek
> "Every time I fire a linguist, the performance of the speech recognizer goes up"
As far as I know, modern, successful NLP systems don't have much human knowledge baked in and are produced by training on large data sets.
In the meantime, here's a preview of what it can do. http://goo.gl/Lz7Vr
> Stub