Thanks! Yup, this is basically custom music embeddings + nearest-neighbor vector search.
I personally found that vector representations performed significantly better than other approaches.
And the results will actually be a lot better once I ship a better model (the current one can definitely be improved upon).