Generating Faces with Deconvolution Networks
zo7.github.io
zo7.github.io
That Kraftwerk/Daft Punk voice, the original idea developed by Bell Labs was to come with ways to compress voice signals maximally by representing them only by the envelope necessary to modulate a bank of oscillators. Like any codec/compression scheme, the basic idea is to come up with a simple set of basis vectors that can represent the same data in a minimal fashion as long as it can be reconstructed on the other end.
http://www.bell-labs.com/newsroom/publications/290623/
"The method specifies the speech signal in terms of its short-time amplitude and phase spectra."
In the case of machine learning, the idea is simply to come up with those basis vectors empirically and automatically instead of reducing them to a mathematical model such as "set of sinusoidal waves".
It's not a huge benefit over what you're suggesting (realistically, when are your bandwidth constraints ever that strict?) but it's an interesting idea.
It's almost too easy, really...