They announced, in the talk, that this is all done locally on the phone. The phone will have a database of, iirc, 10,000 songs locally.
[edit] typos
[edit] typos
I'm probably underestimating the parameters per node, and overestimating the size of the layers closer to the input. Further, it's more likely structured as an LSTM than a convolutional network, since sound is a streaming source.