Industrial-Strength Natural Language Processing in Python
spacy.io
spacy.io
Ironic timing here! We're just preparing the 1.7 release, which has a lot of nice changes, including the option of a much smaller model for English (50mb), to help people test faster.
This means that if you install the library right now, you'll have to redownload the data once the new version is released.
So, maybe wait until tomorrow to get started? Definitely our most ambivalent front-paging yet!
Coincidentally, I wrote a blog post [1] that went up just this morning that, in part, compares spaCy with the other giant in the Python NLP ecosystem, NLTK. TLDR - I think that, right now, the majority of users are better served by spaCy than NLTK.
[1] https://automatedinsights.com/blog/the-python-nlp-ccosystem-...
I'm mainly interested in the Danish language right now, although I might also have a use for Italian so it would be nice to know so as to budget my time.
import spacy
import numpy
from spacy.attrs import LOWER, IS_STOP
nlp = spacy.load('en')
doc = nlp(u'The quick brown fox...')
array = doc.to_array([LOWER, IS_STOP])
content = array[1, numpy.nonzero(array[0])]
Personally I normally work in Cython when it needs to be fast. I find this more productive and more readable than trying to guess what numpy operations will be fast. So I would be doing: cdef void get_tokens(uint64_t* content, Doc doc) nogil:
for i in range(doc.length):
token = &doc.c[i]
if Lexeme.c_check_flag(token.lex.flags, IS_STOP):
content[i] = token.lex.lowerI'm using the CoreNLP C# wrapper, so I'm wondering if something similar (.NET Core compatible) is available/doable for spaCy?
NLTK is better as a learning tool and for messing around. SpaCy is better if you just want something that "just works".