They only need to store (and occasionally update) the kernel parameters for the trained deep neural net. Very small indeed.
I'm probably underestimating the parameters per node, and overestimating the size of the layers closer to the input. Further, it's more likely structured as an LSTM than a convolutional network, since sound is a streaming source.