A Short Chronology of Deep Learning for Tabular Data
sebastianraschka.com
sebastianraschka.com
Even though Deep Learning on Tabular data is a topic that is picking up interest, I'd try to stay away from it as much as possible. The main advantage of "statistical methods" (in contrast to Deep Learning) is the interpretability. Most business applications of Machine Learning happen on "tabular" data (explicit features) with a lot of knowledge by the company about the features selected.
A simple Decision Tree gives you AMAZING accuracy and you can still understand what's going on behind the model.
Interpretability in Business ML applications is, IMO at least, the single most important trait of the model selected.
What is SHAP?
In more detail: the business data is prone to: evolving processes / systems / products / markets / customers; errors / omissions / corrections; tail events; hirings / firings; data loss; etc etc. The datasets tend to be small, messy, complicated, subjective. Nothing about this suggests needing a large, complicated model.
I am curious if there is a sample threshold where it's worth exploring deep learning approaches to tabular data. I wonder if there are other considerations (e.g., inference speed, explainability, etc.).
What was really interesting was when the dataset had more than 6k or so; the deep learning model was suddenly much more accurate and by a wide gap! At roughly the 10k mark, the DL model was outperforming the tree model easily.
Not especially, but there are tasks where DL models seem to occasionally outperform by a little. If you really want to milk extra accuracy it can be worth it to try a DL model, and if it performs as well/better you can use it to make an ensemble along with your GBM or replace the GBM, though it's rare that it is worth it. If you check tabular data kaggle winner writeups most use gbms or an ensemble for a tiny boost over just a GBM.
Assuming limited time to work on the problem, you'd almost always want to focus on further feature engineering first and likely some hyperparameter tuning second.
The major downside of DL is the slow training, and therefore slow iteration feedback loop. Couple that with an exponentially growing number or hparams to tune, and you get something very powerful but costly in terms of time to use.
But if you want the best possible accuracy, and data collection isn't expensive, DL is the way to go. Just expect to spend 10x the amount of time tuning it vs trees to get a 10% to 20% reduction in error.