SpaCy 2.0 released
github.com
github.com
The English neural network named entity model is a huge improvement over the v1 model. However, the training data is still all from 2010, so it makes some notable errors. We're working on improved training data, using our annotation tool Prodigy (https://prodi.gy).
The NER for the other languages is trained on "silver standard" data from Wikipedia, so the quality is much less consistent, especially if you're working with social media text or "chat bot"-type inputs.
We bootstrapped the company by doing consulting, and now we're releasing products adjacent to spaCy. We've had a great response to our annotation tool Prodigy, which is currently in free beta: https://prodi.gy .
The license model for Prodigy is pretty simple: permanent per-seat licenses, with pricing that compares pretty favourably to other developer tools.
We're looking forward to releasing some other offerings alongside spaCy. We don't like to say too much because timelines are tough --- we don't want to release something half-finished to stick to a schedule. The Explosion AI mailing list is the best way to stay in the loop.
Can you be more precise ? Because I don't want to invest time in a tool to discover that I can't afford it months later ( having a research student budget and all, i.e my "budget" is my own money ).
:(
For research students, we think your institution should be covering you! We'll be offering an academic subscription, so research institutions can pay a yearly flat fee to have all staff and students covered.
> How are you this fine morning?
and didn't recognize anything - I was expecting "you" and "(this) morning" to be highlighted - but perhaps I misunderstood the example.
This https://demos.explosion.ai/displacy/ looks really nice though. Just the UX of horizontal scrolling using mouse is horrible (once you figure it out).
Thanks to you and your team for all the hard work!
Curious enough, did u guys train your model on your own dataset or public dataset? Didn't find too much information about the specifics in the documentation
You might find this method particularly useful for meeting memory constraints: https://spacy.io/api/vocab#prune_vectors . This lets you reduce a large word vectors table to a small one by remembering the nearest neighbours for the words you prune out. So if you have a rare word like 'biophysicist', you can map it to the vector for a word like 'scientist', and get a close-enough word vector for it.
We've developed a great product for our customer with SpaCy, it wouldn't be possible without SpaCy.
Can anybody share what kind of projects you're doing that benefit from SpaCy? Do you use it as-is or do you build on top of it?
I haven't checked it out myself yet, so I wanted to ask that are the performance issues fixed that were haunting the 2.0 alpha version?
So if you have an event based system where you can process only a single document at once, it does not make sense to upgrade yet, because for a single document case the runtime performance was 10x-100x slower, at least with 2.0 alpha version.
I'm getting around 8k words per second on the smallest Google Cloud instances. You couldn't run spaCy 1 on these instances (or on AWS lambda) due to memory usage problems, especially problems predicting memory usage for long-running processes. This is why we say spaCy 2 is cheaper to run in a cents-per-word sense than spaCy 1. This is the performance measure that we think is most important.
However, users are still reporting performance problems, so I wouldn't call the issue resolved. spaCy 1 managed to avoid depending on numpy during prediction, making it easy to ensure that performance didn't depend on anyone's environment. spaCy 2 currently does use numpy, introducing these questions around configuration. I'm working to fix this by implementing the forward pass entirely in Cython.
Basically: donations can only be made from personal funds, but most of the benefits from the software will go to commercial users. That's a pretty lopsided dynamic.
My understanding is that there are actually some very good Python libraries for Korean NLP? It's now much easier to provide annotations via another library. This is how the Chinese and Japanese support is working at the moment. We'll add "native" models for all of these languages, but for now you might want to wrap some of these resources: https://github.com/datanada/Awesome-Korean-NLP
https://github.com/crownpku/Awesome-Chinese-NLP
Very much looking forward Chinese support in SpaCy.