What are the downsides of transfer learning? How can it fail?
And do you just arbitrarily select the "cut off output layer" for the pretrained model when retraining with your own data on new layers?
And do you just arbitrarily select the "cut off output layer" for the pretrained model when retraining with your own data on new layers?
Some other areas are much more challenging. For example, in natural language processing tasks you will sometimes see some benefit from using pretrained embeddings, but it is very task and model specific. There's some exciting work going on in this area though.