Ah! Then I'm sure they have some way of making the data needed much smaller than I thought it could be!
I'm probably underestimating the parameters per node, and overestimating the size of the layers closer to the input. Further, it's more likely structured as an LSTM than a convolutional network, since sound is a streaming source.