Is this similar to how Shazam operates? I suppose there's more filtering and denoising going on with mic input.
So my approach is intentionally a bit fuzzier than something like Shazam, since apart from exact matches, you also ideally want a continuous representation for music similarity (so you can find songs that are 80-90% similar, for example).
That is hard to do for approaches optimized at exact matches (which usually use something like audio fingerprinting, etc.)