> While such techniques have been explored before for speech, we are the first to make it work for 48 kHz sampled stereo audio (i.e., CD quality), which is the standard for music distribution.
Reading this extremely charitably, I think the contrast is supposed to be to speech which is normally uses much smaller sample rates. E.g. AMR-WB samples at 16 kHz. Speex, back in the day, supported up to 32 kHz.
Full quality digitally distributed audio is often sampled at 48 kHz (e.g. Opus - the default audio codec on YouTube and many other sources), so I think "CD quality" is just supposed to emphasize that it's full-band rather than wide band or narrow band.