But all in all, that content is relatively rare, and generally transient even in the music they appear in.
But all in all, that content is relatively rare, and generally transient even in the music they appear in.
1. It's a high frequency complex waveform with a fast envelope, so it demands bitrate.
2. Drum miking often involves multiple mics spaced apart, so more than one typically picks up any given cymbal with a phase offset, and those mics are panned quite differently, leading to a very "wide" result, i.e., left and right output is fairly uncorrelated as seen on a vectorscope [0].
3. A perceptual codec at a given total bitrate often sounds better when stored as a mid-side transformation (instead of storing a left channel and a right channel, store a L+R "mid" a.k.a. sum channel and a L-R "side" a.k.a. difference channel), also known as "joint stereo" which is a common flag on MP3 encoders, because it allows for assigning more bits to the mid channel (correlated signals) and fewer bits to the side channel (uncorrelated signals). More bits for mono center-panned stuff like vocals is the goal, which is generally for the best, but fewer bits remain available for wide stuff like those cymbals! Contrast with regular stereo mode where half of the total bitrate is assigned to each channel. MP3 below 256kbps typically needs joint stereo mode enabled in order to sound decent.
If anyone has a good cymbal crash sample at 24/96 or better that they can provide, it seems like it would be a great example for intentional differentiation of various compressed versions.
Grungy rock music might only have a few instruments, but they're often purposefully highly distorted and have people pretty much screaming and shouting, leading to the actual sinal being closer to literally noise.
So the closer you are to literally noise, the less compressible your signal is.
Imagine a an image with a dozen sharp, clear, colorful squares. Now imagine a similar resolution image with only 5 colors, but they're different shapes and they're kind of fuzzy and they're really more like gradients instead of a pure color. Which is going to compress easier?
(FWIW, I’m way more of a punk fan myself, and usually find most classical music pretty boring.)
No, I do agree having multiple instruments does lead to a wider sound than just a single one. Plus multiple instruments will probably help balance out a single one being out of tune or not quite hitting the note right. But still, these instruments usually are way more tuned to produce closer to pure tones and their harmonic overtones than a guitar going through half a dozen different distortion and effect pedals then through a compressor along with a guy screaming all over the place into a microphone.
Also, when strumming a guitar you're almost always playing essentially six strings at once, playing a whole chord with only one instrument. Meanwhile on a flute or a trumpet or a clarinet or a violin a single player is only playing a single note at a time, a single string of the guitar. So during strumming sections a guitar is almost like 6 instruments, in terms of signal complexity. So a rhythm guitar strumming and a lead guitar picking strings is really almost like 7 instruments played by two people compared to many orchestral instruments.
Just look at these two spectrograms. Look at the rock song where there's a lot of distorted guitar, bass, drums, and singing going on and compare that to an active part of the classical recording. See how the classical recording has a lot more clean, straight lines while the rock song is a lot more fuzzy? Imagine if these were images, which would be more complicated to accurately compress? That's not really a great analogy, but it is touching on the same concept.
Rock song: https://youtu.be/BVsp23B8dWo?t=62
Classical song: https://youtu.be/Txp-pHU2K6w?t=210