Why aren’t you using pretrained models?
lrz.me
lrz.me
The most obvious one is that your dataset doesn't share a lot of features with the pretrained dataset. This is pretty common in vision tasks where the pretrained set is usually ImageNet. If there isn't significant mutual information in the tasks (e.g. language to vision or vise versa), then you aren't going to transfer the knowledge that well and you've just wasted your time.
Scientific datasets often have this issue as usually you can't collect much data and even though there might be shared knowledge between the pre-trained it will give you a bias that you don't want or can't use. It may even prevent it from generalizing. Training from scratch can smooth this out in some cases. Pretraining means you're starting from a different point in the optimization space and it can pigeonhole you towards certain optima.
What's often being suggested for pre-trained models are HUGE. Sometimes you might as well just write something from scratch (or train a smaller model from scratch) because you just don't have the compute. There are many small models that are highly powerful and can be trained from scratch. You can even train transformers on CPUs. Various architectures will help with different tasks, so even just a random pretrained model that does well on a dataset isn't going to save you. You may also just be wasting significant/costly compute.
So when you can, use a pretrained model. But knowledge transfer isn't going to always help you. Your millage may vary is all I'm saying and pretrained models are not the solution to everything.
If you've seen a paper that verifies this I'd love to see it. The early layers of the network detecting simple features like lines, curves, textures, shapes, colors, etc could still benefit from what they learn on ImageNet even if the features later in the network are not similar.
FWIW I have not yet seen a model starting from a pre-trained COCO checkpoint that does worse than random initialization.
This is pretty common, especially with convolutions, but is not guaranteed. How you embed matters a lot. For example, there are transformers that use early convolutions for the embeddings and that makes them just work. Though too many and they are less performant (ViT tried pre-resnets which wasn't great). [0] also investigates transfer learning in medical domains and shows that CNNs depends more on statistics reuse and transformers depend more on feature reuse.
> If you've seen a paper that verifies this I'd love to see it.
[0] also discusses this. When it does and doesn't work on medical domains. Basically any paper that discusses transfer learning will also discuss the limitations. But note that there is a bias towards results that work. [1] also shows some of these results, where Imagenet pretraining helps and doesn't (and references others doing the same). Note in Figures 2 and 3 how InceptionV{3,4} and MNASNet have higher performance without pretraining (Fig 4 is a summary). So this shows in part of what I was saying that it isn't always about dataset either. You have a coupled problem that is hard to disentangle. There's also plenty of papers that try to say that LLMs are good at learning vision classification and never get past 50/60% accuracy (or worse) on ImageNet. Lots of scientific papers will also just straight up train from scratch and not mention transfer learning because it just didn't work for them, but you'd need to physically talk to these people as it isn't in their papers.
> FWIW I have not yet seen a model starting from a pre-trained COCO checkpoint that does worse than random initialization.
Additionally ImageNet performance doesn't correlate 1-to-1 with how well it works as a backbone in object detection and segmentation.
As another note, I would often say to be careful with pretrained models. The vast majority of papers are using test accuracy to hyper-parameter tune. So you're leaking knowledge into your model. I think this is mostly caused by reviewer benchmarkism (desk reject if you aren't SOTA) so bad practices become standard.
I ask you kindly to stop showing people, so I can keep feeling smart for knowing discrete mathematics /lh
If it works then great. If it doesn’t, it’s difficult to know why a need even more difficult to fix it. The fix might involve retraining with better data, retraining with different architecture, regularisation, endless and unknown knobs to tune.
Fine tuning the whole thing could potentially take far more work (but could also be worth it - the answer is probably that this depends on the practical use case you've got in mind).
The intuition for this is that lines, colors, textures, shapes are all general concepts that can be learned from a different domain & used in the earlier layers of the model to build up to more complex features.