>I have a neural net onboard my phone which automatically detects songs offline and tells me what they are.
MP3 uses something called psycho acoustics, which is a quantitative model on human perception, which is used to eliminate frequencies that can't be heard based on this model.
Your neural network doesn't tell you what features make songs distinct, it's not a quantitative model at all, but a black box heuristic on what the important features are superficially. If actual mathematicians worked on this problem, I guarantee you they'd do a better job, and their models would work on a commadore64, with real time training. Moreover it would tell you things like who is singing, if it's a live performance, which concert it was.