I wish to know more about the "voice-to-phoneme" compressor he mentions. Anyone have any links handy? Google is failing me.
It sounds to me like you build a library of common sounds that voices make (these are the phonemes), and then you use just combinations of these phonemes at specific times and volumes to get as close to the target waveform as possible. Then you store the sequence of phonemes, which is probably a triplet of integers (phoneme ID, timestamp, volume level), along with a waveform representing the difference between your phoneme-generated waveform and the original. You compress both of them. Since the difference is much smaller than the original waveform, you save lots of space.