My goal was to detect instruments in complex, full-length music, as opposed to single sources with no background noise. My approach was to generate Mel spectrograms for small sections of songs and then run the deep learning classifier on these images to build a list of labels for later use, e.g. generating playlists. I more than doubled the accuracy by adding to and curating my own dataset. I used Resnet50 as the base model and found decent results even when applied to spectrograms!
I still want to do more analysis to figure out what kind of small features the model was actually picking up on. I also want to try experiments like scrambling the spectrogram time-wise and seeing if the results are still as accurate.