FastText – Library for fast text representation and classification
github.com
github.com
Bag of Tricks for Efficient Text Classification: https://arxiv.org/abs/1607.01759v2
Enriching Word Vectors with Subword Information: https://arxiv.org/abs/1607.04606
Both fantastic papers. For those who aren't aware, Mikolov also helped create word2vec.
One curious thing: this seems to use heirarchal softmax instead of the "negative sampling" described in their earlier paper http://arxiv.org/abs/1310.4546, despite that paper reporting that "negative sampling" is more computationally efficient and of similar quality. Anyone know why that might be?
It says this: fastText is a library for efficient learning of word representations and sentence classification.
What does that meant? Is for sentiment analysis?
Sentence classification is the generic term for bucketing sentences into different labels - those labels could be "positive", "negative" and "neutral", thus allowing for sentiment analysis.
But they could also be other labels such as "sports_news" or "finance_news". This library allows both.
This library basically means you don't have to write the code for sentiment analysis anymore (just one example).
Just feed it a model:
$ ./fasttext supervised -input train.txt -output model
And then you can predict what the most likely label for a text is: $ ./fasttext predict model.bin test.txtTraditionally, word representations are learned by looking at surrounding words. So "good" and "bad" will have similar word representations.
By training on sentiment, similar sentiment words should be clustered together.
Adding comments back in would be a great start to contributing to OSS.
Help - how to I format blocks of code/bash output in this editor ?
`fastText josephmisiti$ cat train.tsv | head -n 2 1 1 A series of escapades demonstrating the adage that what is good for the goose is also good for the gander , some of which occasionally amuses but none of which amounts to much of a story . 1 2 1 A series of escapades demonstrating the adage that what is good for the goose 2
Are they saying to reformat it like this
cat train.tsv | head -n 10 | awk -F '\t' '{print "__label__"$4 "\t" $3 }'`
giving me
`fastText josephmisiti$ cat train.tsv | head -n 10 | awk -F '\t' '{print "__label__"$4 "\t" $3 }' __label__1 A series of escapades demonstrating the adage that what is good for the goose is also good for the gander , some of which occasionally amuses but none of which amounts to much of a story . __label__2 A series of escapades demonstrating the adage that what is good for the goose __label__2 A series __label__2 A __label__2 series __label__2 of escapades demonstrating the adage that what is good for the goose __label__2 of __label__2 escapades demonstrating the adage that what is good for the goose __label__2 escapades __label__2 demonstrating the adage that what is good for the goose`
I'm looking forward to try it on some bigger datasets.
I also wonder if is it possible to use separately trained word vector model for the supervised task?
I'm quite surprised that (a) they found classifier4j at all, (b) they bothered to test it, and (c) it won, but anyway...
[0] http://classifier4j.sourceforge.net/
[1] https://www.semanticscholar.org/paper/A-Quantitative-and-Qua...
[2] https://www.semanticscholar.org/paper/A-Quantitative-and-Qua...
My main issue is that the FastText paper [7] only compares to other intensive deep methods and not to comparable performance focused systems like Vowpal Wabbit or BIDMach.
Many of the features implemented in FastText have been existing in Vowpal Wabbit (VW) for many years. Vowpal Wabbit also serves as a test bed for many other interesting, but all highly performant, ideas, and has reasonable strong documentation. The command line interface is highly intuitive and it will burn through your datasets quickly. You can recreate FastText in VW with a few command line options[6].
BIDMach is focused on "rooflining", or working out the exact performance characteristics of the hardware and aiming to maximize those[4]. While VW doesn't have word2vec, BIDMach does[5], and more generally word2vec isn't going to be a major slow point in your systems as word2vec is actually pretty speedy.
To quote from my last comment in [1] regarding features:
Behind the speed of both methods [VW and FastText] is use of ngrams^, the feature hashing trick (think Bloom filter except for features) that has been the basis of VW since it began, hierarchical softmax (think finding an item in O(log n) using a balanced binary tree instead of an O(n) array traversal) and using a shallow instead of deep model.
^ Illustrating ngrams: "the cat sat on the mat" => "the cat", "cat sat", "sat on", "on the", "the mat" - you lose complex positional and ordering information but for many text classification tasks that's fine.
[1]: https://news.ycombinator.com/item?id=12063296
[2]: https://github.com/JohnLangford/vowpal_wabbit
[3]: https://github.com/BIDData/BIDMach
[4]: https://github.com/BIDData/BIDMach/wiki/Benchmarks#Reuters_D...
[5]: https://github.com/BIDData/BIDMach/blob/master/src/main/scal...
[6]: https://twitter.com/haldaume3/status/751208719145328640
I was a bit curious why they did not offer this by default. It seems quite useful.
Here are the almost-comparable evaluations:
fastText Numberbatch
en:RW .46 .601
en:ws353 .73 .802
fr:rg65 .67 .789
The difference actually should be larger: Numberbatch considers missing vocabulary to be a problem, and takes a loss of accuracy accordingly, while FastText just dropped their out-of-vocabulary words and reported them as a separate statistic.I'm using their Table 3 here. I don't know how Table 2 relates, or why their French score goes down with more data in that table.
What's the trick? Prior knowledge, and not expecting one neural net to learn everything. Numberbatch knows a lot of things about a lot of words because of ConceptNet, it knows which words are forms of the same word because it uses a lemmatizer, and it uses distributional information from word2vec and GloVe.
We will add more functionalities for multi label classification in the future (predict the top k labels, etc...).
Deleted comment
Quotes from the paper:
Both char-CNN and VDCNN are trained on a NVIDIA Tesla K40 GPU, while our models are trained on a CPU using 20 threads.
Table2 shows that methods using convolutions are several orders of magnitude slower than fastText.
Our speed-up compared to CNN based methods increases with the size of the dataset, going up to atleast a 15, 000× speed-up.
Table 2 shows the speedups of:
ConvNets: 2 to 5 days on GPUs
FastText: 52 seconds on CPU