You are reading it wrong. The things they are patenting look like the following:
1. A neural network system implemented by one or more computers, wherein the neural network system is configured to generate an output sequence of audio data that comprises a respective audio sample at each of a plurality of time steps, and wherein the neural network system comprises:
- a convolutional subnetwork comprising one or more audio-processing convolutional neural network layers, wherein the convolutional subnetwork is configured to, for each of the plurality of time steps:
- receive a current sequence of audio data that comprises the respective audio sample at each time step that precedes the time step in the output sequence, and
- process the current sequence of audio data to generate an alternative representation for the time step; and
- an output layer, wherein the output layer is configured to, for each of the plurality of time steps:
- receive the alternative representation for the time step, and
- process the alternative representation for the time step to generate an output that defines a score distribution over a plurality of possible audio samples for the time step.
This appears to me to be describing an autoencoder for audio compression.
There are two other claims as well, all related specifically to audio processing.
Now whether those actual claims are valid is a separate question. I don't know the state of the art in 2016, they could have been. They are still math, I agree. But that's how you tell what they are actually claiming, you can entirely skip the body of the patent and just read the numbered claims.
The ones that don't start with "The ... of Claim ..., wherein" are independent claims meaning that they stand on their own. The ones that do read that way are typically used to narrow the independent claim so that when they say the independent claim is too broad they can make those dependent claims mandatory.
So for instance, if the USPTO (probably) rightly claims that generating audio with a convolutional neural network isn't patentable, they can fall back to saying "The neural network system of claim 1, wherein the audio-processing convolutional neural network layers include one or more dilated convolutional neural network layers." is valid combined with claim 1 because no one has thought of that specifically, and perhaps that's all that is granted.
In that case, as long as you don't use a dilated convolutional neural network layer then it is not infringing.
It's still all ridiculous nonsense, but that's how you read the patent.