What Kagglers Are Using for Text Classification
mlwhiz.com
mlwhiz.com
Kaggle prioritizes chasing a metric, but real-world data science has more considerations.
Chasing a single metric is certainly too narrow but getting the best results often does matter in a professional context too. While Kaggle can go overboard on massive ensembles using a state-of-the-art approach to the problem is often warranted outside of Kaggle.
In my experience, the worst text classification is usually fine. Labels are usually too inaccurate and subjective for "accuracy" to matter much.
Besides that, the accuracy gains are not marginal anymore (BoW can't compete like it used to, especially with pre-trained models).
Consider for instance an RSS reader that classifies articles to determine whether or not to interrupt the user with a notification. This should be fast to train and update the model on the fly every time the user enters a correction (e.g. 'this article actually isn't interesting', or 'interrupt me with articles like this in the future'.)
If you are deploying on resource-constrainted devices (IE: low-end PC's without GPU), it is not unusual to take a lot of time training a model on a very powerful computer (which nobody cares about), then distilling or transfering the result for test time.
This isn't true. It depends on your priorities and goals. Machine learning that spends most of its time unable to learn is not real AI. Some of us are interested in sample and energy efficient learning capable of on-line incremental updates immune to catastrophic forgetting. Not just because this is truer to actual learning but because it moves away from being dependent on a handful of companies to do the actual training.
Anticipating some replies: no, transfer learning or meta-learning methods don't really avoid this. In the case of transfer learning, you still have that high coupling between a handful of sources. The down-sides of this is its own discussion. In addition, there are times where the ability to extract local relations can be dulled by the dominant wikipedia and common-crawl representations. Meta-learning gets you fast updates but you still cannot stray too far away from the domains that were met at training time.
> What matters is prediction speeds
I'm not a fan of bag of words models either but a simple dot product is always going to be faster than many matrix multiplies and or convolutions. The implementor should always try these as a base-line and decide if the performance accuracy trade-off is worth it for them.
Online learning, sample - and energy efficiency are unrelated to training times. Like said: nobody cares if you ran Vowpal Wabbit for 1 hour or 100 hours, as long as you are not constantly babysitting it and calling that paid work (or have the unusual requirement of daily retraining while using an online model).
> simple dot product is always going to be faster than many matrix multiplies
If you care about this (because it is profitable), you rewrite in lower-level language or predict with cloud GPU (which will be at least comparable to simple dot product, while adding performance)
A parallel discussion we are having is whether the gain in accuracy is always worth the gain in complexity and loss in speed. It's something to decide on a case by case basis. It's basic hygiene to reach for the simplest model first.
What is proper AI? It's all dumb curve fitting right now.
LOTS of people care how long it takes to train a model. A few minutes, vs. a day, vs. a week, vs. a month? Yea, that matters.
Think about how long it takes to try out different hyperparameters or make other adjustments while conducting research...
If you're Google maybe you don't care as much because you can fire off a hundred different jobs at once, but if you're a resource-limited mere mortal, yea, that wait time adds up.
If we are talking days or hours: start parameter search on Friday and return best parameters on Monday.
Do research and iteration on heavily subsampled datasets.
If you are building models for yourself, or for Kaggle, you may care in as much as your laptop gets uncomfortably hot.
Another important aspect is training and incremental training on edge device.
At the time when privacy is becoming very important and you cannot export data from mobile devices etc. Training time on mobile is an important factor
I do consider the cloud both widely available and near infinite in resource adding capability.
If it is really not economically feasible to add resources, then the performance gains were not as promising as thought (whether cloud or on-site).
1) The ML experts in the field have all, pretty much, settled on the need for a uniform method to train models, but for each model needing to be trained on-site.
2) While the cloud might be near infinite in terms of adding capacity, "Hey guys, lets stage up some health-data compatible AWS instances to do something that was a side project we're not even sure will work" in what is always a cash-starved part of healthcare is...well...a pretty big ask.
In the future, you could use “most”.
That's a reckless generalization. I care.
My thesis would take forever if I didn't do any optimization. Also my data is 20 rows with ~6000 predictors.
There are models out there that can take months! I worked on one that took months. We had to tweak it and optimize it to see if we can get it to acceptable training time.
In kaggle some competitions it takes over 7 hours to train a model, and I can generally think of 10 things a day to try. prediction only takes about a minute.
> "especially with pre-trained models" if the corpus are different, pre-trained models do not help much, if not hurt.
I dont use NN because they simply don't have great accuracy, and most importantly they have a huge amount of variance. this is mostly because the data on kaggle is not very large. the gbm trifecta (xgboost, catboost, lgbm) also does really really well.
> "Kaggle prioritizes chasing a metric, but real-world data science has more considerations."
this is counter to your point. most real-world considerations need things like model explanability.
I notice that things dont make hacker news if they are doing anything other than NN. model with xgboost gets X% accuracy - crickets. model with X-Y% accuracy with DNN - headline news.
I also notice that teams in industry tend to throw a DNN at a problem and never try something more simpler like xgboost. I saw a team with an LSTM for text lament they had 80% accuracy on training/evaluation, but when pushed to prod dropped down to 50%. I saw the errors they were getting, I said: maybe its too complicated and not generalizing well, have you tried xgboost?
they retorted LSTM's w/ word2vec is very robust. I thought, obviously its not given your results. I tried to offer the idea that word2vec was trained on an entirely different kind of corpus (also a problem when trying to use word2vec in kaggle)
However, what's lacking in the ML practitioner community is nuance. Some applications need deep models some problems need xgboost. There isn't a "best" model in text classification because it depends on your data and problem.
https://scikit-learn.org/stable/tutorial/machine_learning_ma...
dimentionality reduction is part of any step, which depends on number of observations, features, and model. my favorite is l1 or pca, but I am not afraid to use stepwise regression or some tree method.
Let's say I know seasonality is a strong feature in classifying my text, how can I add this? With a BOW I can literally just add SEASON_AUTUMN as a word to the text, and I'll get that extra feature as a dummy variable in my feature vector.
But for fasttext, if I add such a word, it will just be averaged out in the final document feature vector.
That might not be true as it might increase bias and thus might need a more careful hyperparameter tuning to avoid overfitting.
That blog used a 2d cnn because tensorflow didn't have a 1d version at the time of writing, so he just created a dummy 2nd dimension of length 1 and called it a day.
It's resume-driven-development for data scientists.
I've never seen an interviewer impressed with the fact that a job was performed using not-deep learning, but say that you used deep learning (despite how spurious it might be) and they light up like it's Christmas.
In my own tests on my own corpuses, CPU-based FastText is faster to train and produces significantly better results (precision/recall) than the GPU-bound CNN algorithms that I've tried, but have not compared it against RNN techniques.
(Not for winning kaggle but for an actual problem)
It's also worth considering that you might be best off going with none of these options. Cool as deep learning is, I've personally never actually been able to justify using it in a professional setting. Simpler models such as logistic regression and decision trees have characteristics that are near-useless for getting you to the top of a Kaggle leaderboard, but can be indispensable when working on many real-world business problems"
- anonymous comment reply
Attention was first coined in this paper (as far as I know): https://arxiv.org/pdf/1409.0473v7.pdf
The second page of the introduction of the "Hierarchical Attention Networks for Document Classification" paper mentioned in the article even cites it.
Here's a summary from the end of each section...
1. TextCNN: "This kernel scored around 0.661 on the public leaderboard."
2. BiDirectional RNN: 0.671
3. Attention Models: 0.682
I remember getting 95%-ish accuracy with BoW and the SVM circa 2004 when it came to questions like "is this paper about astrophysics or organic chemistry?"
In that case you have a distinct vocabulary for different topics and it is hard to beat BoW.
Sentiment analysis, on the other hand, is where BoW goes to die since now "not good" means something very different than "good", and even simple heuristics like treating "not X" as a term that is different from "X" give limited gain because negation is expressed with constructions like "i don't believe that is good" and there is no k-word window that you reliably catch negation in since there isn't a limit on how complex sentences are.
There is also the question of "is the improvement between method A and method B worth it?" For instance the Netflix prize was much celebrated because some brilliant people busted their ass to go from 92% to 95% accuracy on movie recommendations. In the end the algorithm proved to be too complex for the value it created. (eg. Who would notice that they got 8 bad recommendations instead of 5 out of a hundred? An additional half a bad recommendation out of 10?)
The real "Netflix optimization problem" is how to spend as little on acquiring content as possible while motivating people to keep their subscriptions and that is something Netflix will keep closer to their chest and not promote a public competition on. (eg. if it were valuable why would they let competitors know about it?)
Nothing inherently wrong with these methods, just a lot of possibilities for misuse in my eyes.