SpaCy: Industrial-strength NLP with Python and Cython
honnibal.github.io
honnibal.github.io
The Cython implementation makes it somewhat believable that it's faster than CoreNLP, but I'd also like to hear a deep-dive on why it's several times faster, beyond that control over memory layout is the best way to win performance (stipulated). In particular, it would be good to know whether CoreNLP is doing more processing than spaCy or otherwise handling more concerns.
Finally, I'd really love to see a feature table comparing spaCy with CoreNLP.
Compelling work!
Time complexity. The Stanford parser is a phrase structure parser that creates dependencies as a post processing step. So, assuming that they use some variation of CKY, the time complexity is O(N^3 |G|) where |G| is the size of the grammar. This uses Nivre-style greedy parsing, which is O(N).
So, a slightly fairer comparison would be e.g. the Malt parser. Although this will probably be better than that in terms of accuracy, since last time I checked the Malt parser doesn't use dynamic oracles yet and by default doesn't integrate Brown clusters or word embeddings (though you could do that yourself). Though I wonder a bit about feature set construction, because in my experience perceptrons are far more sensitive to adding 'wrong' features than e.g. SVM classifiers. This becomes interesting especially when you train a model for another language or dependency scheme annotation, since the features that are relevant differ per set-up.
First of all, the target output of the two systems is exactly the same --- labelled dependency parses with the same schemes, parts-of-speech, tokens, and lemmas. At least, that's how I run the CoreNLP in this benchmark. It has some other processing modules, but I turn them off for the speed comparison.
Second, very very similar algorithms are being run. The new CoreNLP model uses greedy shift-reduce dependency parsing, same as spaCy. That CoreNLP model was published late last year; before that, CoreNLP only had the older polynomial-time parsing algorithms implemented, which are much slower, and often less accurate.
You can see their paper here:
http://cs.stanford.edu/~danqi/papers/emnlp2014.pdf
The contribution of Chen and Manning's paper is to use a neural network model, where I'm using a linear model. (More specifically: they show some interesting tricks to make the neural network actually perform well. I suspect many people have tried to do this and failed.)
Chen and Manning say that their model is much faster than a linear model, because the linear model must explicitly compute lots of conjunction features --- I use about 100 feature templates.
So, they probably have something of an algorithmic advantage over my parser, although the extent of it is unclear. I'll only know when I implement their model. It's not terribly hard to do --- it's just a neural network --- but it's lower on my queue than a number of other things I want to work on. My hunch is that I won't see nearly as much benefit from it as their results suggest, because their baseline is quite weak.
So, I do think all we're seeing here is the same algorithm implemented in Java and C, so the C version is coming out 7x quicker. This makes sense to me. But, possibly the CoreNLP parser has to do some contortions to integrate into their framework. I don't know.
There's also a meta-level point. Maybe I just tried harder. The Stanford paper would still have been accepted, and still have been great, if it ran at 50% of the speed that it does. And we'll probably never know what would happen if the author spent a month doing nothing but trying to optimise the code --- I can't imagine he/she ever will. That wouldn't get a publication.
For your other question, about what spaCy offers and what CoreNLP offers. These are the main things I'm missing at the moment:
* Named entity recognition
* Phrase-structure parsing
* Coreference resolution
I have some preliminary work on NER. I plan to roll that out next, along with some word-sense disambiguation. PSG parsing is no problem to do either.
Thanks for the suggestion to include an evaluation of OpenNLP -- I'll do that.
Also, it would be great if you include SENNA into your benchmark: http://ml.nec-labs.com/senna/
I think the other interesting problem to tackle currently is training data. The situation for English is ok, if you have a couple of thousand of dollars to spare for a commercial license (which may be problematic for bootstrapping). But for many other languages there aren't even treebanks available that can be used for commercial purposes.
It would be great if some annotation project started that aimed to provide annotations under a liberal license.
(Ps. I have a statistical dependency parser written in Go, which I will probably release soon in case anyone is interested ;).)
I agree that the data situation is troubling, though. I don't understand why Google gave the English Web Treebank to the LDC. Why not just distribute it themselves?
The LDC is really more of a problem than a help now. For instance, the OntoNotes corpus costs $0 for non-commercial use. Great! How do you get it? Send the LDC a fax, and when they get around to it, they send you a log in to their ancient website.
It used to be a valuable service to host and distribute this data. Now, this is no longer really the case, but it's still standard to distribute via them.
I haven't read LDC's license on Penn Treebank recently, but AFAIR you cannot just redistribute models that were trained on the Penn Treebank. Or put differently, you can distribute the model, but any users still have to obtain a license for the treebank. That's why we are still stuck with the Brown corpus and such.
I don't understand why Google gave the English Web Treebank to the LDC. Why not just distribute it themselves?
Indeed.
Just checked the license and my memory seems to be correct:
https://catalog.ldc.upenn.edu/license/ldc-non-members-agreem...
:(
I do hate the proliferation of the idea of BSD-style licenses being "commercially friendly". You can do one hell of a lot of things commercially with GPL software. It's only under very specific circumstances that you are obliged to release source for GPL software. And as (I would wager probably) ~99% of NLP software is used internal to a company and never distributed as such, the differentiation between licensing models would never actually come into play.
I'm curious why you think there's a market failure here. If a library like this can produce N salaries of value, then it should be able to earn N salaries of revenue. Maybe it takes >N salaries of work to produce it --- in which case, okay. That means this project is more costly than it is useful.
I definitely believe that market failures exist, and are quite common. But the service I'm trying to sell is trying to create economic value very directly. I'm not writing poetry here.
Say a patch/pull-request is freely offered to your AGPL project. I believe it's well-understood that such a submission includes implicit permission for distribution under that same open-source license.
But you will also be licensing that same contributed code under another custom commercial license. Contributors should understand that – you've made it clear in surrounding docs – but whether you can legally assume their re-licensing permission, without an explicit copyright-assignment, is a bit murkier of an issue.
Aside from that, the market failure I foresee for your project is the following: Say you keep building this out and write a fantastic, state of the art general purpose NLP python library. Now an a academic like myself comes along and forks the AGPL version of your repository, contributing additional functionality to the parts of the pipeline I am most familiar with. You cannot re-incorporate my work into your commercial license (unless I sign those rights away, which I won't), so now you're stuck trying to license an inferior version of my fork. Meanwhile, since mine truly is just open source, my version can freely accept both bug fixes and added functionality from one-off contributors who are using said fork. Better yet, unless you change your business model, I can continue to re-fold in any changes you make upstream, as well as include parts of other GPL libraries that build up in the intervening time.
Now, I'm not saying this is a perfect argument. Perhaps enough people are still interested in paying you for a a commercial version of the original software, but I think in the long run as the two version diverge that's unlikely to generate sufficient revenue for you.
I'd also note that you gain absolutely nothing from maintaining your separate fork, other than the principle of the thing. I'm compelled to distribute the code under the AGPL, just as you are. If your features are compelling you could instead negotiate with me for a cut of the license fee.
The best way as far as I'm aware of to get into NLP is to take a course at a university. My university (University of Melbourne) was lucky to have an undergraduate subject taught by one of the lead authors of NLTK (Stephen Bird) and that was a great help. You can even take the subject on its own without enrolling in a full course.
Plus there are books on it. One of the main ones I'm aware of (by Bird) is http://www.chegg.com/textbooks/natural-language-processing-w...
A Python 3 edition will be released next year.
Pattern doesn't really use machine learning, just some pre-computed statistics from the annotated data, and some hand-crafted rules. Machine learning is good. It's really the right way to build these systems.
Sorry you're dealing with so many licensing questions here but a quick clarification:
> If their company is acquired, the license will be transferred to the company acquiring them. However, to use spaCy in another product, they will have to buy a second license.
Is the second license only required because they sold the company on (and the license along with it), or is a license per product generally required? In other words, if I buy a single license, can I make and sell two different products?
One license allows you to develop one product. If you stop work on one thing you can re-use the license on something else, though --- it would be silly to ask at what point a change of focus becomes a different product.
This seemed the sanest way to do it. I think per-site, per-user etc licenses are really stupid. The license then impinges on your technical decisions.
I've thought a lot about passive reduced relative clauses over the years --- they were a big part of my PhD thesis. So I happen to know that the first one in the WSJ data is wsj_0003.1. This isn't in the training or development data, but it's in the same data set --- so, this is a fair but optimistic spot-check. The sentence is:
> A form of asbestos once used to make Kent cigarette filters has caused a high percentage of cancer deaths among a group of workers exposed to it more than 30 years ago, researchers reported.
There are two reduced passives here --- "used" and "exposed", and a potential (but unlikely) false positive in "reported".
spaCy correctly attached "exposed" to "workers", but didn't attach "used" correctly --- it attached it to "reported" instead of "form". This doesn't really make syntactic sense, but that's what it did --- the system's entirely statistical; there's no grammar licensing certain attachments.
To see the parse, run:
from spacy.en import English
nlp = English()
tokens = nlp(u'A form of asbestos once used to make Kent cigarette filters has caused a high percentage of cancer deaths among a group of workers exposed to it more than 30 years ago, researchers reported.')
for word in tokens:
print word.orth_, word.tag_, word.dep_, word.head.orth_[1] http://alt.qcri.org/semeval2014/task8 (with 2015 rerun)
edit: noticed OP offers trial license at $1: http://honnibal.github.io/spaCy/license.html
I appreciate the thought, but if this isn't useful to support itself, then obviously I was wrong, and I should find a more valuable project. But I don't think that's the case --- I think this will help a lot of people build useful products, so the commercial license should fund its development quite adequately.
If you want to increase sale, you could include donors' names in the documentation in return to $100, for example.
(Except obviously not.)
spaCy's POS tagger works like the one in the blog post, but it's implemented in Cython, and has some extra features.
Pattern has some nice morphological processing features. I don't do morphological generation, for instance, and I haven't hooked up the morphological analysis to the Python API yet.
TextBlob also gives you a few extra bits and pieces, like a wrapper of the Google translate API.
I'm really targeting the situation where you want to build a product around some NLP. In this use-case, you need the NLP to be fast, you need it to be as accurate as anyone knows how to make such a system, and you need it to be entirely in your control.
As far as GenSim goes: it's good. It does different things from spaCy, though --- topic modelling, etc. It would be nice to interoperate between the two libraries. I have no plans to implement topic models.
I'm a fan of "do one thing, do it well".
Having said that, it would be great to facilitate "spaCy + gensim" pipelines for users.
For example, the "word vector representations" can be trained easily with gensim, on arbitrary user-specified corpora, whereas spaCy loads something pre-trained, in a specific format. Maybe room for some interoperability there?
I'm not sure why this author is further propagating FUD that suggests GPL code is unsuitable for commercial use. Just because companies are irrationally afraid of the GPL doesn't make it true.
Or are you saying companies should be more willing to GPL their products? That's unlikely to happen.
My understanding is that if you link to the library, your code must also be GPL, which means that anyone linking to your code must be GPL, etc.
This is a problem if you're trying to sell your code. Probably you don't want to make it GPL, and you probably don't want to force your customers to make their code GPL.
I think relatively few tech companies are trying to sell their code directly these days. Most are hoping to build a product and/or service and then charge customers for access to it.