Prodigy: A new tool for radically efficient machine teaching
explosion.ai
explosion.ai
Fei-Fei Li gives a good sense for this in her history of ImageNet [1][2].
[1] https://qz.com/1034972/the-data-that-changed-the-direction-o...
[2] http://image-net.org/challenges/talks_2017/imagenet_ilsvrc20...
1. This seems closer to a reinforcement learning system than a pure annotation system. That seems to be by design, however based on the demo, I am not able to change or add to the annotations as I go, which is a big limitation. It's just yes, no (no feedback), ignore and undo. This is in contrast to something like the VGG annotations system: http://www.robots.ox.ac.uk/~vgg/software/via/via.html
2. I don't see an actual annotations capability for images in the demo. Not sure if that is just a pretotype page, but IMO image classification/segmentation is the place where this tool would really benefit the community.
3. It's unclear to me how or if I retrieve my trained model or even just the annotated structure (.csv?, .json?) from this system. Do I get a .pb somehow that I can import into TF or am I locked into an API with my new model served from Prodigy? My guess would be the latter.
I think what this wants to be is a human validation system for training, which also improves the Prodigy nets through crowd sourcing. Definitely a win-win in the short term, but it has the limitations of the initial model and the ability for the user/client to tweak the system and output the results.
Matroid is doing something similar here, but I have been unimpressed with their offering so far.
For the specific questions:
1. The built-in web views all have binary annotation interfaces. This is more of a design choice than a fundamental limitation, and the front-end is extensible --- you can add your own web views if you need to.
The binary interface is sort of a position statement. We think this is The Way, so we want you to try it. We'll have more input components in future, but at the start we want to guide people towards the intended workflow.
2. The beta focuses on NLP support, but there's a front-end for image classification, and a workflow page: https://prodi.gy/docs/workflow-image-classification .
3. You can usually get some accuracy improvement by retraining once all the annotations are available. I've not found a streaming SGD algorithm that works as well as the simple iterate-and-shuffle batch process. Batch training also lets you tune the hyper-paramers. You can read more about this here: https://prodi.gy/docs/workflow-named-entity-recognition#trai...
I would suggest writing a Prodigy recipe to do the batch training. That way you can pass in the dataset ID, instead of exporting the annotations. There's no problem with exporting the annotations and running a script, though. Again --- it's all on your computer. You can run it however you like.
You're right, I totally missed that. Re-reading that comes through but it is definitely different than what I would have expected.
Thanks for the response, I'll dig in further.
Now coming back to the topic, I have so far just used Jupyter Notebooks and spreadsheets to do annotations and by golly, it is an extremely boring and tedious process. This looks like a fun tool to try out for my next NLP related project. Might spice things up!
But I hope that like all SpaCy related ideas, it doesn't assume too much about the problem at hand. I usually use NLTK instead of SpaCy because it allows me to be very flexible, except for the sentence tokenizer, where SpaCy's accuracy is hard to beat.
For those looking for alternative OSS solutions: BRAT, labellmg are decent.
Sounds like marketing BS. what about OpenNLP and Stanford's for NLP?
Agreed, the description is definitely cringe-worthy.
As if whoever wrote that wasn't aware that these are language geeks they're marketing to.
In other words: Spacy is sinking.
It's just that, wording-wise "the leading open source X" exudes marketing-speak, which I find language geeks tend to have robust antibodies against.
This kind of lingo work (sort of) for the market, say, MongoDB is in. But for the users of these tools, I suspect not so much.
* only supports 3 languages, though
Stanford CoreNLP give good accuracy and is pretty much the benchmark in English for accuracy. BUT it isn't great software. It falls over if you pass large amounts of text to it, the code is dreadful, it's hard to integrate (even in Java because of its own wacky config system), various parts aren't integrated (eg, SUTime), it doesn't have an embedding representation and it is pretty slow.
Having said all that I still use it sometimes. But Spacy is much nicer to use, and 99% (probably more) of the time the slightly lower accuracy is offset by things like the easy availability of word embedding right with the word tags.
I think it's pretty fair to say Spacy is the leading open-source NLP tool.
Getting labelled data is a pain.
(But I'm definitely not an expert)
Great work here, btw. It's refreshing to see emphasis on "label some damned data" and work towards making that easy.
Are the examples picked those that have the highest objective function error rate, or something similar?
Does this apply only to text classification problems? Are there examples where this could be applied to tabular data?
They kind of have this weird dissing of unsupervised scenarios, though. It's not like supervised or unsupervised is better or worse, they're just surrounding different problems. They can talk up their product without needing to criticize a problem domain.
It's like if you were making motors for boats, and then started talking about "these crazy people who think it's better to fly." ???
I do think there's a pretty common failure mode for teams who don't have much experience with ML, though. Teams who don't have much experience with ML often take "We don't have much data" as a parameter of their problem, and don't see that this is something they can decide to change. This can lead to a lot of time spent experimenting with different unsupervised approaches that are a poor fit for what they're trying to do.
https://github.com/tensorflow/models/blob/master/syntaxnet/g...
We're mostly worried about saying we "support" a language when we've just trained a tagger on a UD treebank, though. We like at least having the stop words and tokenizer exceptions filled in by a native speaker, so the usual flow has been that someone needs the functionality, and they make a pull request.
If you just need the UD model for say, Bulgarian, you can do:
python -m spacy train xx /path/to/output_model /path/to/bulgarian-train.conllu /path/to/bulgarian-dev.conllu --no-entities
We don't have a spacy.bg.Bulgarian language class yet, so you can either add one, or use the multi-language class, which usually works OK.Sorry about the poor performance on the site! We got complacent because all of our sites are 100% static.
It looks like it is online during typing of annotation, trying to predict annotations.
When it says teaching, it means teaching the AI. When it says "radical" it means ... getting slightly more data input, and in an online manner.
Examples for text data: topic tags, marking mentions of companies, finding descriptions of protein interactions. For images people mostly do segmentation (finding object boundaries) and classification.
If the model is generating images or text, you usually need to do annotation to evaluate the model. For instance, you need to know whether the translations are grammatical, whether an image is coloured correctly, etc.
patio11 blog is HTTP only but this blog is HTTPS.
HTTPS without keepalive is likely to kill any cheap VPS, establishing HTTPS connection is intensive.
That being said, the core of the issue is that they should use nginx (or apache in mpm-events). And they should have cloudflare.
More intensive than adding numbers together, sure, but computers are pretty fast. If you're doing 100k connections a second you might have to give some thought to that. Meanwhile, if you have KeepAlive on, 2~5 clients per second will kill you.