Try out Stanford's CoreNLP natural language software
corenlp.run
corenlp.run
But.. there are problems. As software engineers, the (many) authors make great researchers.
CoreNLP is wonderful in the many different ways it almost lets you integrate into it without hacking the code. Changing the config of the various components is fantastic, because you get a very comprehensive view of many different ways people can configure almost the same thing. Environment variables? System properties? Properties files in a specific location? In the classpath? Json config? YAML? It almost supports them all - or rather different parts use different ones, and the only way to work out exactly how it works is to read the code.
Also, the licensing is annoying. Everyone doing commercial stuff with it just puts it behind a web service anyway, so they should just embrace non-viral license and get some input from the community.
Also SUTime. Yes, it works mostly, but wow :(
(Sorry for the rantish post. I use parts of CoreNLP a lot, and I'd love to see it improve.)
From memory when I was last looking around I cared mostly about named entity recognition (NER) and Spacy themselves say CoreNLP is better. CoreNLP has more features too.
WordVector integration in Spacy looks interesting though. That's probably enough to make me have a play with it.
Spacy is really fast, the author is extremely knowledgeable, and it works well on the datasets it was trained on. Problem with SpaCy for me was that it was pre-trained on those texts and it was not possible to train it on new things. Also their parser wasn't very customizable.
This pre-trained model is borderline useless if you want to obtain good results on your data, which is probably very different from the data they trained on.
For NER in the Python world, the best option is pycrfsuite. It works really quickly and lets you easily define your own features. CRFSuite itself is a work of art. I only wish Okazaki incorporated 2nd order transitions because that makes a huge difference on some datasets.
CoreNLP is a lot worse than pyCRFsuite in terms of ease of integration and performance in my experience. Particularly, if you want to define your own features.
CRFSuite
I've never used this, but it's not really a ready-to-use NLP toolkit is it? Isn't it more a tool for building NLP tools with?
If you want to do NER in a way that doesn't suck there is no way around making your own model on your own training data.
It honestly takes only a few days of labeling things yourself. I found that outsourcing the work to amazon turk is not a viable option because the graders there are terrible. And they work about 30x slower than you do. Even if you pay them $1/hr, that is like paying one person $30/hr. I'm not kidding.
Sure you can do a quick and dirty "send data to these guys and they'll do all the work", but I haven't come across a model that works well on all datasets. We're talking going from 30ish percent accuracy for a model not trained on your dataset to low 90s for something trained on your dataset. Of course, these are approximate numbers and it is definitely possible that your dataset is almost exactly like the ones they trained their model on.
It's incredibly simple to make your own model.
1. Label your data with brat: http://brat.nlplab.org/index.html # 5 days for 2k one page documents.
2. Tokenize data with nltk/spaCy. Come up with features and label using pycrfsuite: http://nbviewer.ipython.org/github/tpeng/python-crfsuite/blo... # 1 day
3. Do more labeling, retokenizing, neural embedding from word2vec's similar words to the tokens you have, tune parameters or come up with better features such as your own dictionaries of entities, etc. Retrain the model. #2 weeks.
4.Done. Now you have a memory efficient fast model tuned on your data. You can label anything you want. Not just Person/Company, but things like car vs bicycle brands, computer parts, obfuscated email addresses, etc.
I want to make "domain adaptation as a service" the key part of spaCy's business model: you send us text, we send you a good model. Internally this will probably involve annotating part of the text, but that's a tactical decision we'll make.
I hope we can make some break-throughs that help NER be much more general than it is currently. But the current solution you describe works fine; it's just a pain in the ass for each organization to take on. We want to have the required infrastructure and expertise set up, and make the process seamless.
I has recently changed to MIT.
Also SUTime. Yes, it works mostly, but wow :(
So what's better than SUTime? I'm using it now and am genuinely curious.But the code and config mechanisms are terrible to use and the documentation of them is even worse.
It's ok if you want to use it as it is out of the box. But try to do something like change to forward looking dates ("On Monday" should mean next Monday instead of last Monday) and it isn't as easy as it should be!
Originally I got interested because of an essay PG wrote about Bayesnian spam filtering, so I wrote a Bayesian classifier in Java (this was over 10 years ago, when that was pretty cutting edge).
That led to text summarisation - still open source Java stuff, and apparently now considered state-of-the-art[1] (I'm kind of amazed, because that code is 10 years old).
Then I did some AdTech stuff, wrote an open domain natural language question answering thing (like Watson, but not as good - but it was just for fun. bAIb is the way to approach this now if anyone is interested).
Now-days I'm using NLP for future event prediction.
See https://research.facebook.com/researchers/1543934539189348 and http://smerity.com/articles/2015/keras_qa.html
I see at least that identifies most words as FW, "foreign word", perhaps?
I'm not very familiar with CoreNLP (or NLP, more generally), does it only support english out of the box? Or is it nothing out of the box and this is configured in english?
Google translate does a pretty good job of guessing the language of non-ridiculously-short-and-ambiguous sentences. Also, from what I know about how it works, it seems quite agnostic to specific languages. Does a less stochastic approach (as what I assume NLP does) provide such flexibility? Something akin to the nlp library knowing all languages, and deciding on which of them a given sentence makes sense. I can't put an example on the table right now, but surely there are sentences that can work on more than one language; at least when you accept a couple of misspells.
I remember reading about computers being confused by this phrase in 80s and 90s. Apparently, not much progress here.
Which humans with only one pass can't either.
You can download the spaCy library for yourself if you suspect I've cooked the example :)
http://spacy.io/displacy/?full=fruit%20flies%20like%20a%20ba...
What is it getting wrong?
The correct parse has "eating" as an xcomp of "likes", and "sausage" as a dobj of "eating". You can see the correct parse for this structure if you plug "I like eating sausage" into the demo.
EDIT: Weirdly, the demo of the same software at http://nlp.stanford.edu:8080/parser/ doesn't make this mistake!
Weirdly, the demo of the same software at http://nlp.stanford.edu:8080/parser/ doesn't make this mistake!
See my rant about the config of CoreNLP: https://news.ycombinator.com/item?id=10350090
You're right that the configuration is confusing for new users. But, you have to remember that this is first and foremost research code intended to be flexible enough for the Stanford NLP group's research needs. In terms of particular configuration sources, nearly everything should be configurable from properties passed in as a properties file. Are there exceptions to this rule?
SUTime is documented to accept a property to a file that is read to obtain other properties. Quote:
sutime.rules = [path to rules file][1]
I'm unclear if that works - looking at our code it appears the properties files need to be in the classpath under a package that is defined by the sutime.rules property.
I don't remember how other packages worked.
Didn't get it right.
(See https://en.wikipedia.org/wiki/Buffalo_buffalo_Buffalo_buffal... )
Did better on "Colorless green ideas sleep furiously."
The real problem is that the parser comes so swiftly to wrong conclusion and cheerfully presents it as a valid result.
It would look a lot better if it simply reported "error: cannot parse that". (Better yet, with reasons: "I cannot parse that because I get stuck on this specific ambiguity and it's just too much for me.").
Also, what about the possibility of multiple results? Language is ambiguous. If something has two parses, it's wrong to assert just one.
This thing has made no consideration whatsoever that even a single instance of "buffalo" in the sentence might conceivably be a verb, which flies in the face of almost any noun in English being verbable.
It's almost purely syntactic reasoning. Searching these spaces of possibilities is something which, you would think, a "natural language parser" ought to be doing to earn its name.
Nobody actually knows what it means "to buffalo" something; it is not necessary to know. People solve the parse in spite of knowing that there is nothing to understand in the sentence.