NLP concepts with spaCy tutorial
gist.github.com
gist.github.com
It is truly groundbreaking, and a major improvement over NLTK. I also recommend gensim, another phenomenal library for NLP.
SpaCy really fills an important gap
Oh, I see now that it's quite a lot further along then when I last checked in. ~$400 for an individual license seems pretty fair IMO.
We're really happy with how Prodigy's being received. It's only been on sale two months, so I'm looking forward to hearing more success stories as people finish their projects (and of course, feedback to change what needs to be changed!).
You can read how FullFact used it to train claim identification models for fact-checking here: https://fullfact.org/blog/2018/feb/how-we-customised-prodigy...
Probably the best place to follow the progress is the support forum: https://support.prodi.gy/
We're also working on more tutorial videos. This video shows the workflow for training a new entity type: https://www.youtube.com/watch?v=l4scwf8KeIA . This is one of the bits of the tool we're particularly proud of --- you can start off with a couple of seed terms, use word vectors to build up a larger terminology list, and then turn that list into a set of pattern rules to start boot-strapping a classifier. Prodigy will suggest phrases that match the patterns as entities, and your answers are used to train the statistical model. As you keep annotating, the model will start suggesting phrases too, which you'll say yes or no to. Eventually the model basically takes over, and you're mostly correcting its suggestions.
https://github.com/explosion/thinc#no-computational-graph--j...
https://arxiv.org/abs/1711.10455
Could someone give me some more detail?
About these neural networks, I think it's "just" implementation. Here's the linear layer implementation in Chainer: https://github.com/chainer/chainer/blob/master/chainer/funct...
We have the forward and backward pass organized as class methods here, and the intermediate state from the forward pass is saved into attributes in the instance. So on each call to the network, we make an instance of this LinearFunction class.
In terms of what's being computed, there's really no difference between this and what happens when you call a layer in Thinc. It's just that the state gets captured in the outer scope of the closure. Maybe Thinc's way has a little less overhead, if there are fewer levels of indirection. Thinc uses the Chainer folks' GPU library --- so, unsurprisingly if you define the same network, the benchmarks are very similar.
On the other hand...I do think the implementation matters! Here's a difference for you: if the library approaches it as "we're going to build a computational graph, and execute it", then the library is going to steal the control flow. If the library tells you "here are some functions, and some higher order functions to compose them", you have more access. PyTorch and Chainer doesn't steal the control flow to nearly the extent that Tensorflow does, but they still build up and tear down the state in their objects, and that makes it harder to intrude.
(I'm the author of spaCy and Thinc)
There is no such thing as a sentence, or a phrase, or a part
of speech, or even a "word"---these are all pareidolic
fantasies occasioned by glints of sunlight we see reflected
on the surface of the ocean of language; fantasies that we
comfort ourselves with when faced with language's infinite
and unknowable variability.
'Pareidolic' is my new favourite word.Linking this to signal theory and the Fourier transform, one point to consider is that solutions are only true in the infinite limit, so a word, a phrase is never enough to represent reality. A sense of continuity is real enough, but discontinuity, too, although I can't position that in a psychological frame. Or Neurological. But speaking with the y combinator in mind, I don't think words are the fixpoints of thought, but feelings are. Maybe onomatopoetic names are and familiar faces are close enough.
https://www.semanticscholar.org/search?q=named%20entity%20li...
This survey paper I coauthored is outdated now, but it's not the worst place to start: https://www.semanticscholar.org/paper/Evaluating-Entity-Link...
In practice it's very driven by information retrieval, especially the coverage of the synonym list provided in the ontology. Relevance scores for each ontology entry are also super important.
Disambiguation is getting better with neural networks, but it's still really hard. If you have two entities matching for some text, picking the most frequent one gives you a very strong baseline.
We've evaluated and used quite a few over the years: there's Dexter [0] by Diego Ceccarelli, Semanticizer by UvA [1] and DBpedia Spotlight [2] and a few others. We've used them for various linking tasks, such as detecting "work skills" in plain text (HR domain) or detecting drug names (medical domain).
The amount to which these tools allow "customization" (ease of plugging in your own ontology, support for input format and disambiguation signals) differs. Either way, even though this research includes open source code, the code is more of the "research prototype" kind. Don't expect a plug&play optimized production tool.
[0] http://dexter.isti.cnr.it/
I am wondering how people solve actual problems with spacy. What are the use cases? Is Spacy used in question answering sytems or summarization pipelines? Maybe conceptual search?
I prefer the approach of link grammar / relex which are based on dictionaries / grammars, it seems easier and less error prone.
Prodigy is genius! More tools like that will be built in the next few years...
Can you link to something open source that outperforms spacy on the basic NLP tasks (NER, POS, dependency parsing)?