It's even worse if the sound is going through a digital (VOIP / cell phone) system. Most modern codecs for voice are using some form of Linear Predictive Coding[1] (e.g. ACELPC) which is basically modeling sound as a resonator at the bottom of a tube with a filter bank (sortof like your voice box). With voice, this is a reasonably good approximation, and the codecs are aggressively tuned to be efficient at that. But if full-band music gets piped through it will sound roughly like that music is being produced by a flapping plosive at the bottom of a long tube.