Winograd transform can be seen as a depthwise convolution with fixed weights.
And similarly to DTF, winograd transform can be used to speed up convolutions: a 3x3 convolution can be implemented with a winograd transform + a 1x1 convolution + inverse winograd transform. This basically reduce the number of operations performed by 2.
Now, if a neural network is able to learn the weights of a depthwise convolution to match the coefficients of a winograd transform, that means you should be able to do this transformation when describing the architecture of your network. That way, you don't rely on your deep learning framework to do the transformation for you.
And more importantly you might end up with something more generic: perhaps the weights in the depthwise convolutions will not converge to the weights of a winograd transform, but to weight better suited to what you are trying to learn.
This is the same ideas which lead to the development of CNN: replace the fixed weights of convolutions used in traditional computer vision algorithm with weights that are learned by the network.
The property of matrix multiplications is that they are composable, i.e. `X * (Y * z) = (X * Y) * z` that is, in the end you only need one matrix.
So what this means in practice is that you have FFT for free. NN is doing a matrix multiplication anyway. Discrete time Fourier Transform is a matrix multiplication. Thus it can simply fold DTFT together with whatever other transform it is doing - it doesn't cost anything.