Shazam works with a signal processing technique ("fingerprinting" - https://patents.justia.com/assignee/shazam-entertainment-ltd) that aims to re-recognize the song as originally recorded - so its design goal was "Which song is playing here?" or more precisely "Which song from my database is played here (again)?" rather than "Which is the closest match from my database to this new song rendition/'cover'?".
You can imagine it like a sequence of hash codes or shingles (for modeling gaps/pauses to borrow a term from Web page similarity) for subsequent parts of the songs.
Notably, Shazam does not aim to transcribe the lyrics; so the OP's approach may potentially claim some novelty here. In any case, this experiment shows how great large pre-trained neural language models are for rapid prototyping to put something together quickly - perhaps to test feasibility before attempting to develop something better and more bespoke.
In a related note, Apple may be working on a similar service, at least they filed a related patent application: https://patents.google.com/patent/US6990453B2/en