Scaling Down Deep Learning
greydanus.github.io
greydanus.github.io
I'm surprised we haven't seen more research on modular models. Random forests and boosting both work by building a collection of weak models, and combining their results to produce final predictions. It seems like building an ensemble of some sort of unsupervised learners, then training a smaller supervised network on that ensemble could produce good results, while also being more reusable and easier to train.
In my opinion the right approach to downsizing is to train at full size and then either distill or prune the model to keep the parts that are relevant to your problem.
Unfortuntaly, distillation of smaller models can take a fair bit of time. However, there is a lot of recent work to make distillation more efficient, e.g. by not just training on the label distributions of of a teacher model, but by also learning to emulate the teacher's attention, hidden layer outputs, etc.
More parameters is inherently better, with the knowledge we have now.
If your concerns are about over fitting there are lots of regularization techniques used in practice like dropout, weight decay, and data augmentation.
There's nothing preventing you from sharing weights across layers, and would be interesting to see some research about that.
E.g. the ALBERT model does that:
https://arxiv.org/abs/1909.11942
I have done model distillation of XLM-RoBERTa into ALBERT-based models with multiple layer groups and for the tasks that I was working on (syntax) it works really well.
E.g. we have gone from a finetuned ~1000MiB XLM-R base model to a 74MiB ALBERT-based model with barely any loss in accuracy.
Open access to the paper here: https://doi.org/10.1093/gigascience/giaa119
For example, we are in the process of writing a paper for the compression of proteins. There are a few papers about this problem, but only one working compressor that we could find. Even more problematic is when the papers in question claim very good efficiency and we are always left wondering if we are missing something, or if the data they used isn't the same data we have access to etc. In that regard, GigaScience is great, because it places a lot of importance in reproducibility.
Also, one can train the network to a good accuracy, then change the input layer, and unfreeze the inner layers, that way the network will have a head start.
Not sure how universal this principle is, but it seemed reasonable, if I remember it correctly, of course.
The approach described in the article looks very smart. Also could be handy for integration testing of ML frameworks. I've been working on my own DL framework, and this data set looks like a good way to test the training and inference pipelines E2E.
1. 10 classes is way too small to make meaningful estimates as to how well a model will do on a "proper" image datatset such as COCO or ImageNet. Even a vastly more complicated dataset like CIFAR-10 does not hold up.
2. I feel like CIFAR-100 is widely used as the dataset that you envision MNIST-1D to be. Personally I found that some training methods will work very well on CIFAR-100 but not that well on ImageNet so TinyImageNet is now my go-to "verify new ideas dataset"
I guess I'm just asking if COCO or ImageNet-trained networks are actually noticeably superior for most real-world tasks, or if it's just a metric that's used because the performance differences only show up in the long tail of the distribution.
Given that for any real-world vision task you start from a pretrained model om those datasets they will in fact be noticably superior on the real world task after finetuning. Just because the quality of the features extracted through the backbone is better.
Or on Medium: https://medium.com/deep-learning-reviews/scaling-down-deep-l...
How would I make it easier for people like you to read?
Also, my blog has kinda unique colors. Do u find it easy to read?