MP3 vs. AAC vs. FLAC vs. CD (2008)
stereophile.com
stereophile.com
- Q: Can you hear the difference between CD-quality lossless audio and anything higher fidelity? A: No, no one even has the biological ability to. 44100khz, 16-bit audio can perfectly reproduce audio as far as we can physically tell. The only reason to store anything higher is for production or archiving (that is, for computers to listen to).
- Q: Can you hear the difference between 320kbps MP3 or the equivalent, and CD-quality lossless? A: Yes, this is _theoretically_ possible. However, many well-controlled listening tests have been performed on this subject that all say no, so it's much more likely that you can't, and the burden is on you to prove otherwise with an abundance of evidence.
e.g.: https://downloads.bbc.co.uk/rd/pubs/whp/whp-pdf-files/WHP384... https://www.researchgate.net/publication/257068576_Subjectiv...
The listening test linked in the article leads nowhere, I would have liked to see their methodology.
But all in all, that content is relatively rare, and generally transient even in the music they appear in.
1. It's a high frequency complex waveform with a fast envelope, so it demands bitrate.
2. Drum miking often involves multiple mics spaced apart, so more than one typically picks up any given cymbal with a phase offset, and those mics are panned quite differently, leading to a very "wide" result, i.e., left and right output is fairly uncorrelated as seen on a vectorscope [0].
3. A perceptual codec at a given total bitrate often sounds better when stored as a mid-side transformation (instead of storing a left channel and a right channel, store a L+R "mid" a.k.a. sum channel and a L-R "side" a.k.a. difference channel), also known as "joint stereo" which is a common flag on MP3 encoders, because it allows for assigning more bits to the mid channel (correlated signals) and fewer bits to the side channel (uncorrelated signals). More bits for mono center-panned stuff like vocals is the goal, which is generally for the best, but fewer bits remain available for wide stuff like those cymbals! Contrast with regular stereo mode where half of the total bitrate is assigned to each channel. MP3 below 256kbps typically needs joint stereo mode enabled in order to sound decent.
If anyone has a good cymbal crash sample at 24/96 or better that they can provide, it seems like it would be a great example for intentional differentiation of various compressed versions.
Grungy rock music might only have a few instruments, but they're often purposefully highly distorted and have people pretty much screaming and shouting, leading to the actual sinal being closer to literally noise.
So the closer you are to literally noise, the less compressible your signal is.
Imagine a an image with a dozen sharp, clear, colorful squares. Now imagine a similar resolution image with only 5 colors, but they're different shapes and they're kind of fuzzy and they're really more like gradients instead of a pure color. Which is going to compress easier?
(FWIW, I’m way more of a punk fan myself, and usually find most classical music pretty boring.)
No, I do agree having multiple instruments does lead to a wider sound than just a single one. Plus multiple instruments will probably help balance out a single one being out of tune or not quite hitting the note right. But still, these instruments usually are way more tuned to produce closer to pure tones and their harmonic overtones than a guitar going through half a dozen different distortion and effect pedals then through a compressor along with a guy screaming all over the place into a microphone.
Also, when strumming a guitar you're almost always playing essentially six strings at once, playing a whole chord with only one instrument. Meanwhile on a flute or a trumpet or a clarinet or a violin a single player is only playing a single note at a time, a single string of the guitar. So during strumming sections a guitar is almost like 6 instruments, in terms of signal complexity. So a rhythm guitar strumming and a lead guitar picking strings is really almost like 7 instruments played by two people compared to many orchestral instruments.
Just look at these two spectrograms. Look at the rock song where there's a lot of distorted guitar, bass, drums, and singing going on and compare that to an active part of the classical recording. See how the classical recording has a lot more clean, straight lines while the rock song is a lot more fuzzy? Imagine if these were images, which would be more complicated to accurately compress? That's not really a great analogy, but it is touching on the same concept.
Rock song: https://youtu.be/BVsp23B8dWo?t=62
Classical song: https://youtu.be/Txp-pHU2K6w?t=210
128k MP3s, though, fall apart with more complex instrumentation.
It will allow you to choose a codec for streaming that is different than storage.
Easy transcodes to various lossy formats are the primary reason for a well maintained and curated library of proper FLACs (or ALACs).
The main advantage being that your can fit more great sounding music on your space-limited device. Disk space is cheap, but upgrades on your phone are probably not. Even if your Music collection fits easily, why not have more room for other media / apps?
I've heard the resonable assertion that the most gifted audio engineers in the world cannot distinguish 192 kHz sample rates from a raw line feed, but some can distinguish 96 kHz. I certainly can't. I used to build audio equipment. There were legendary "golden ears" that people would drive hours to meet, for design feedback. Whatever they heard was reproducible blind with other "golden ears".
How does this square with the logical assertion that there's a sharp cutoff to our biological ability to hear frequencies? Those tests don't account for our ability to sense the presence/absence of overtones.
Now, computers may want better resolution, not to "listen" so much as to better transform in novel ways. Just as "DogTV" recolorizes for its audience, a computer could make the inaudible audible in novel ways. Reconstructing better 3D sound stages. Accurately reconstructing a singer's facial expressions via AI video, rather than simply splatting out something plausible. The whole point of computers are to extend our reach, coevolving in every conceivable way.
Overtones are frequencies. If you filter the higher frequencies you remove the higher overtones. The only way you're going to "hear" ultrasound is if it's very loud and you hear audio-frequency distortion generated in your own ear, e.g. like when bats are squeaking nearby. But this isn't useful musically, because everybody's ears distort differently.
Are people asserting that an ear removed from a cadaver, hooked up to the best available scientific equipment, measures as a perfect biologically derived low pass filter? Or that we even partially understand how neurons work, when there may be quantum effects to be uncovered a century from now?
Intellectual history is a graveyard of models confused with reality.
This matches my own experience well: most of my friends do not care about various levels of compression, nor what headphones they use - that's fine, I'm glad they're enjoying art in their own way - but I, and some others, do in fact stand to benefit from less compressed audio.
I've personally done blind tests on myself using a python script that randomly plays compressed and uncompressed snippets of the same track and mp3@320 was not transparent to me (though opus@256 was).
Can I tell the difference when casually listening? I don't know, but when the cost of lossless is having my music collection take 60gb instead of 20gb on my 512+gb device, I have no reason not to go for lossless.
If you would give high quality audio experience, to a person that has been listening through 80s general store headphones, to low quality radio rips on magnetic tapes, you might be surprised how few people are going to describe one as "better", without prior description of work and technology required to produce each experience.
And one would be even more surprised by how many people choose the cassette tapes because of nostalgia and a long time satisfying experience.
If loudness was artificially boosted, I had a harder time but could still often tell. I think the sound engineering of the music played a big role and a lot of modern music isn't mixed with complexity in mind.
I also go for lossless ;-).
The audible difference can be described as sound stage size, instrument separation and atmosphere.
The problem is making these details audible needs a good system. “Good” doesn’t mean $10K+ here. Two high quality, large-ish two way bookshelf speakers and a good amp with enough punch (50W+ Yamaha or similar will do) plus a good source in a sizable room is enough.
There’ll be people who can’t tell any difference, there’ll be people who can “feel” it, and there’ll be people who can pinpoint differences. This is because the ear training and biological limits of said people.
I have a friend who can pinpoint a half note (natural vs. sharp) mistake in a 90+ people symphony from YouTube recordings, incl. the instrument. His natural ear resolution is around 1/9th of the tone. He always tunes his instruments via ear and verifies with a tuner. So, this is not impossible.
My ears are not that absolute, but I can divide music to layers and pinpoint details, for example.
Lastly, taking a “diff” of CD quality and 320kbps MP3 version of the same track will leave an audible residue.
There are other comments I left over the years here. Search them for more info. I’m on mobile. I have no practical way to link all of them.
Like for example, the fact there is an "audible" diff is meaningless. The threshold of hearing is not linear nor frequency independent. This is called "masking" and it's exploited by lossy codecs to allow for better encoding as well as audio watermarking. You can add a noise that would be perceptible by itself to content that is entirely masked by the content itself. And the reverse is true, you can remove content without it being perceptible.
The idea of lossy codecs is they filter out the things you theoretically don't hear, yes. However, the presumption that you don't hear these when they are present is not completely true. Because they have a secondary order effects in overall sound.
The audible residue you claim that I don't hear when it's in the CD is the part of the sound which adds this instrument separation and soundstage expansion. Same is for higher sampling rates. While you can't pinpoint the difference with words, it shows itself as smoothness and "richer" sound.
Saying that you can't hear that difference is akin to saying "Human eye can't see faster than 30/60/X FPS anyways", which is not true.
When anyone presented with a lossy-encoded audio file produced with a state of the art encoder and not brick-wall mastered, will be impressed, yes. This includes me, too. However whenever you listen to the same file in lossless or, if present, higher resolution formats, with a sufficiently transparent audio system a couple of times, you start to notice the differences.
There are a couple of caveats in all of this audio business. First of all, you need to know how your audio system sounds and behaves to be able to discern differences. This requires time with the same system for a long time, to understand how it responds. In my case, I have the luck of having the same amplifier (An AKAI AM-2850) for ~30 years. I know how that thing responds to any genre of music, and I know how anything should sound at any quality level. Again, as I aforementioned, you need to do these ABX tests a couple of times back to back, esp. if you don't know the track, to be able to decode the details in sufficient manner. Digitalfeed's ABX test (https://abx.digitalfeed.net) understands this and makes you listen to the same thing 5-10 times according to your available time.
See, I'm an ex-orchestra player. I played in concerts, listened master recordings, and YouTube uploads of the concerts I played as well. I have also listened tons of CDs, MP3s of the same albums, etc. Some of the albums I listen have a captivating sound when I listen to them from CDs. MP3 versions of the same albums do not nail me to my chair, yet I can't leave the CD version of the same album to get a cup of tea. Both are ran through a Yamaha CD-S300 CD player with an iPod interface and MP3 playing capability over USB.
I can also write how CD tracking quality affects audio clarity, but this comment is long enough. In short, Yamaha's old CD-Recorder, CRW-F1 really improved sound quality by abusing Red Book standard by lengthening the pits of audio CDs. It reduced to capacity to 68 minutes, but it was worth it, esp. on lower end CD players.
Sorry this one part especially makes no sense. Digital is digital. Either it added more samples per second, or more bits per sample, or it's snake oil. There's a stream of bits that comes out of the reader. There's no residual information about the length of the pits.
EDIT, yeah sorry this is completely and utterly impossible that you are getting better "audio clarity":
"Yamaha tries to attract computer enabled audiophiles with the Audio Master technology. Audio Master promises reduced jitter and decreased error rates for audio recordings via extended pit and gap sizes on the CD-R. This is actually quite simply achieved by increasing the disc rotation speed vs. the laser clock frequency. In other words, Audio Master recording at 8x rotates the disc at 8.2x, thus creating the extended pit & gap lengths. This naturally reduces the capacity of the disc."
Literally they are just spinning the disc faster, reducing capacity and make it slightly less likely that errors will be read. If you're getting read errors on playback, that means your disc is dirty or your CD player sucks. It's the same bitstream, just read at a different linear rate.
If you honestly believe that this is an audiophile concern, I'd urge you to reevaluate a lot of your other beliefs, because they are clearly not all grounded in technical facts.
DAC's digital part is easy. What differs in quality is the analog part. If DACs were that simple, a 25 cent DAC would power every unit from bottom bin to top tier.
Before that Yamaha CD player I had, I used a lower end Sony CD-Player (I don't remember the model, sorry). Writing the same album, to same brand of CD-R, with the same speed in two different modes created two audibly different disks.
I sometimes challenged myself by writing in both modes, not marking the CD-Rs, and the audio difference was always audible. Even after weeks. 68 minute CDs were always had larger sound stages with more clarity and instrument separation. This is again on the same AKAI AM-2850 amplifier.
I guess this difference would be impossible to hear today, because higher end units have better tracking and better DACs. Also some of them use DAE and use multi-second buffers, so the "slower stream" is no longer present in the pipeline due to buffering.
There's no "slower bitstream" for the DAC. That's provably nonsense and you can work it out from basic principals. The same bits would come out of the optical interface of a CD player, at the same rate either way. If the CD player has a built-in DAC, the same bits would get fed to that same DAC either way.
I'm sorry, but if this is truly what you believe, it really puts everything else that you said into question.
To give you the benefit of the doubt, I might say that the lenses or lasers on your CD players are filthy, and you're just hearing skips or noise from poor reads and that a slower-written, borderline-spec disk might just allow them the function better. Perhaps your player was interpolating or concealing frames [1] that it couldn't read correctly and failed to correct via ECC and you were just hearing a poorly reconstructed digital data stream.
This sort of confident incorrectness, ignoring the underlying technical architecture, is probably why people don't believe anything that an audiophile says.
[1] https://www.pearl-hifi.com/06_Lit_Archive/02_PEARL_Arch/Vol_...
CDs don't record 0s as pits and 1s as lands. A change from pit to land or land to pit is a one, and no change over a time base is a 0.
Therefore, the recording and tracking performance can be affected by the disc content and processing applied.
CDRInfo's tests back in the day showed dramatic improvements in C1 levels, see [0]. Considering some of the lower end CD players by leading manufacturers didn't even had 16bit audio decoding and used late stage oversampling, reducing C1 errors was/is a big deal in recorded media.
As I said in my earlier comment, this mode lead to clearly audible improvements in my older, low-end Sony CD player. I don't how how will it fare in my new Yamaha player due to technology improvements.
[0]: https://www.cdrinfo.com/d7/content/yamaha-crw-f1e-cd-rw?page...
Assuming you have a properly functioning CD player, these errors have zero input in the quality of a CD being played. If you have a noticeable difference in audio quality between a CD with a max error rate of 24 or 32 C1 errors, you've got a tremendously faulty CD player that is complete trash.
Its funny too because in this table it shows the Mitsubishi media performed about identical or better in every speed while being significantly faster at recording. The 1x speed with AudioMaster on for Plasmon is the worst result, ignoring the time they burned 16x media at 44x speed. This table is also challenging to actually compare, because they show different speeds for the different modes (1,4,8x for AM, 4,16,44x for regular) so the only one we can really compare fairly is the 4x. Even then, at 16x speed AM off it had a lower average error rate than 8x speed with it on!
Don't get me wrong, burning a lower error rate from the get go is good, it implies the burn will possibly be more reliable over time as you get things like scratches and other imperfections on the disc. But arguing that a disc with an average of 2.1 C1s vs 1.1 is going to be noticeably different in the sound is absurd. And once again, even then this showed an improvement only in one of the two medias tested. Maybe its better in more media, maybe its worse, maybe there was just something odd with their burns and this is largely just noise in the overall results of burn results from this drive.
I'm still not convinced AM actually did anything but reduce your recording time. Your link doesn't say anything of actually increasing the audio fidelity in the slightest, just that in one straight comparison the C1 errors were lower. Practically every result in that table other than the 16x media being written at 44x speeds is already massively in the "negligible levels." Being below 220 is "negligible", and all of these burns (aside from the one using AM at 1x speed!) are well below that.
However, I'd hope that anyone that cares about fidelity has a CD player that does a little more to generate a DAC clock. NCO run by a software PLL or a hardware PLL with a good loop filter are things I've heard of, but control systems is not my specialty.
I have tried multiple times to discern flac vs 320 mp3 across genres. Every time I believe I can figure it out and I consistently fail to exceed 50% (pure chance) accuracy.
Makes me wonder what ultra-linear source gear or speakers would highlight the differences in real-world situations, if at all. But for my purposes I’ll happily accept the roughly 80% file size reduction for no audible difference.
Brain is interested in the low hanging fruit, i.e. the music and the melody itself, first. The music needs to became mundane or ordinary to be able to listen it deeper for more details. This is when differences can be heard more easily.
Lastly, you don't need perfect systems to hear differences, but understand how your systems respond to the music you're listening to. i.e., your music system's sound needs to be mundane to your brain too to be able to go from low hanging fruit to minute differences you were not able to hear before.
A couple of years later I found an ambient track that sounded messy at 320 Kbps, and that converted me to using flac instead. Disk space had got cheaper enough over that period that it made no meaningful difference.
THis was all years ago now (I'm in my 40s...), and I don't worry about it too much any more. Firstly, I still use flac, so it's identical to the original anyway. And secondly, even if I did use mp3, my aged ears probably couldn't tell the difference even if I turned the volume up to unreasonable levels.
I agree on the kHz (as well as on MP3), but I deeply disagree on 16 bits.
Because yes, if you keep your headphone volume at a single reference level and never turn it up, then 16 bits is fine. This is very much proven.
BUT this ignores the fact that people often turn up the volume a ton to hear the quiet part of the classical music, or on that YouTube video where the volume is inexplicably 5% as loud as it should be.
So in practice, 24-bit audio allows you to retain perfect fidelity even when you have to turn the volume up. 16-bit doesn't.
I don't understand why nobody ever talks about this. (Or why you have to install special utilities on your Mac to be able to turn up the volume to 200% or 400% in order to listen to those YouTube videos that are maddeningly recorded at 5% volume.)
Unless you mean that 24-bit allows for representing audio that is stored at an extremely quiet level at the peaks, wasting most of the dynamic range. That would make more sense - but if audio is printed in such a flawed way, I would expect other quality issues to be present as well.
This is consistent with my other comment about badly encoded MP3 being far from transparent.
Doing the math on the analog bits, the engineering data indicates the analog noise will always end up greater than the 16-bit floor.
"One day" I built an analog preamp which had lower noise than a CD could reproduce anyway.
In any case, while 130dB is a bit excessive, a highly efficient speaker (e.g. the old Voice Of The Theater) connected to a good modern amplifier and DAC could easily be cranked to 118dB with a noise floor that’s, at least in principle, inaudible in any normal room.
(I’m not saying this is a good idea. But seriously, check out the performance of the top amplifiers at audiosciencereview.com — these things have ridiculous performance and aren’t even that expensive. About 120dB SNR at over 100W is something you can just buy, for about $1500.)
(I’m also not claiming anything about linearity of the system or of people’s ears. But I can imagine 16 bits being put to better use in a well-considered floating point system than as plain linear PCM.)
In fact, the upper limit of ~16kHz is defined by the intersection of the “threshold of pain” power curve and the “threshold of hearing” curve. So the human ear has zero dB of dynamic range at the upper frequency limit.
None of this is relevant to a real, dedicated music playback system that doesn't contain a digital mixer. You can't hear noise at -96dB. Your amplifier will swamp that with it's own internal noise sources. In the 80s the audiophools loved to complain that CDs were too quiet because their beloved LP noise was supposed to be desirable for some whack reason.
There was an even a creepy ad campaign several years ago that took advantage of this. They had a billboard in New York for A&Es new show "Paranormal State" with the tagline "It's not your imagination".
They used an ultrasonic system on the billboard to make audible sounds appear in a small region on the sidewalk but not anywhere else. When people walking along the sidewalk got to that region they would hear a woman whisper "Who's there? Who's there? It's not your imagination".
That system worked by making a single ultrasonic beam that somehow as it dispersed became audible. There are other systems that use multiple ultrasonic beams that produce audible sound via interference where the beams meet.
Many acoustic instruments do produce significant amounts of sound above normal human hearing range. Cymbals for example have nearly 70% of the sound power above 20 kHz. Trumpets with a mute have almost 2% above 20 kHz.
It seems possible then that if you wanted to produce a recording that reproduces the sound you would get if live acoustic instruments were playing in the same environment you might need to include ultrasonics unless you are making a binaural recording.
This does raise the question of what we actually want playback of a recording to achieve. Is a recording of a string quartet when played back in my living room supposed to sound like that string quartet is playing in my living room, or is it supposed to sound like what I'd have heard if I was there when the piece was recorded, or is it supposed to be something else?
(For those who haven't heard of binaural recordings, they are stereo recordings made by placing microphones inside the ears of a model human head so they record the sounds that actually ends up in each ear when something is recorded live for a listener at a specific location in an environment. This page of headphone tests [2] includes a binaural test if you'd like to such a recording).
You're not experiencing these ultrasonics with headphones. You're not going to faithfully recreate some ultrasonic interference pattern in a random room with a pair of tower speakers placed in any arrangement. If anything, you're more likely to just make extra noise that wouldn't have been there originally from those interference patterns happening haphazardly and chaotically, coupled with the fact the vast majority of audio gear isn't designed to reproduce ultrasoincs (they're often specifically made to try and not generate them!)
And even then, by recording in the room with the instrument, you are capturing the resulting ultrasonic interference that you could have experienced if you were there.
Compression kills the high end, and learning to recognize tell-tale compression artifacts will forever ruin your ability to appreciate streamed music, low-bandwidth wireless audio systems, or just 320kbps rips of music, certain genres faring worse than others.
I know that in grade school, part of the requirements for joining band, due to the size of my school and overwhelming demand, was passing an audiometry test, where we were evaluated on a few different contexts related to ability to discern detail in audio, such as pitch and volume. I remember being pulled away into the principal's office where some of the test administrators were present, and they accused me of cheating and demanded to know how I did it.
Apparently, I was the only student in the entire state to get a perfect score on that test, at least for that particular year. Unsure if they were implying I was the first ever, but that seems ridiculous to me because passing the test boiled down to just paying close attention.
So I really don't know. Maybe the average person can't hear it, but I know just what to look for in the high-end and usually guess even 320kbps mp3 correctly from my own self-tests wherein I would randomly select between different encodings of a music file. I'm confident I would do well in an administered ABX test if I'm simply being tasked with finding the difference between mp3 lossy encodings and a lossless reference.
I had an experience I mentioned elsewhere in this discussion, about hearing a difference between 24bit 96 kHz and CD-quality in a sound studio. I don't know whether the operator screwed up or what. It doesn't mean I don't think CD quality is "good enough". But that's not the same question as "non-discernable".
On the other end, I have had a few encounters with satellite radio (SiriusXM) and I cannot understand how that product exists. In each car I've encountered it, I've had a visceral reaction and wanted it turned off after a few minutes of trying to listen to it. It put me somewhere between anhedonia and dysphoria.
There are a ton of people in the audiophile world who swear that knowledge of what you’re listening to makes a physiological difference and therefore it’s not the placebo affect. It’s insane, but then so is most audiophile marketing.
You shouldn't conflate that with things like telling the difference between a 44.1KHz and 96KHz sample rate, which is bollocks.
That said, it's a good practice for any archivist to source the highest quality digital versions of albums possible, in case they end up befriending or needing to barter with aliens with much more sensitive ears.
The question is can you hear those differences, to which the answer is basically no, for a modern codec at high enough bitrate. "enough" is 128kbit or more in blind listening tests.
No one here is arguing that modern, high-bitrate codecs aren't much better at producing imperceptible artifacts. But 128kbps is absolutely not enough in blind listening tests, 128kpbs typically produces a ton of perceptible artifacts. You're just making that up.
To a “trained” ear that is listening out for the differences? Sure, those are perceptible. But those aren’t normal people.
AIUI there are some things that don't go away at any bit-rate, e.g. pre-echo.
It may be that two different people have different results of being able to tell the difference due to physical differences, no matter the "effort" the put into listening, or training.
But most of the time that's intentional, as high frequencies you "can't hear" just waste bits that could be used on something more useful, or even cause noise from aliasing and harmonics, as no actual playback equipment is perfect. And extending the frequency range makes it much harder to design the circuitry. All for something you can't hear :P
It's another reason why high frequency (96khz) playback is kinda useless, or even make things worse, as the extra frequency range that gives cannot be heard by humans anyway, and just gives more opportunity for those higher frequency patterns to cause distortion. It may even be that people can "tell the difference" precisely because of those distortions. But that doesn't mean it's "better", indented, or even get the same result on different playback equipment.
Imagine encoding a sort of real world dynamic range across 16-bits. This would go from 0db to 100db in volume. This would need more than 16-bits which yields an SNR of about 96db. The dB values are different and not comparable but you can see we don’t capture the full dynamic range of human hearing very well.
How many encodings does it take before a trained listener using good equipment in an ideal setting can tell?
It's true that, technically, you'll get better results from the second codec when starting from the uncompressed source. Generally, it's always better to avoid unnecessary generation loss. That doesn't necessarily mean that you'll hear a difference since that depends on the cumulative output quality.
A few years ago I did lots of AB testing with some Sony xm1000w3s (Sony LDAC) and Tidal Hifi with some 24bit masters and it was an incredible experience that changed my mind in the whole "640K.. 16bit is enough" argument.
"have some anecdotal evidence"
Also totally fair, I shouldn't be hastily writing comments but I am interested in audio and wanted to share my experiences, naively.
I have some 'hires' recordings that I can't tell the difference between a CD at all and do not care for, and some where I hear more details (on very high end), more separation between instruments - and from what I'm reading it seems more like this is a mastering issue. The difference on some of these recordings enable a kind of subjective 'holographic' spatial effect in me (perhaps the cause of my emphatic response) and it seems I have probably falsely attributed this to the higher resolution as the factor.
http://archimago.blogspot.com/2014/06/24-bit-vs-16-bit-audio...
>In a naturalistic survey of 140 respondents using high quality musical samples sourced from high-resolution 24/96 digital audio collected over 2 months, there was no evidence that 24-bit audio could be appreciably differentiated from the same music dithered down to 16-bits using a basic algorithm (Adobe Audition 3, flat triangular dither, 0.5 bits).
>Furthermore, analysis of those utilizing more expensive audio systems ($6,000+) did not show any evidence of the respondents being able to identify the 24-bit audio. Those using headphones likewise did not show any stronger preference for the higher bit-depth sample. No difference was noted in the "older" (51+ years) age group data (not surprising if there is no discernible difference even with potential age-related hearing acuity changes).
> the effective dynamic range of 16 bit audio reaches 120dB in practice [13], more than fifteen times deeper than the 96dB claim.
> 120dB is greater than the difference between a deserted 'soundproof' room and a sound loud enough to cause hearing damage in seconds.
> 16 bits is enough to store all we can hear, and will be enough forever.
https://people.xiph.org/~xiphmont/demo/neil-young.html#toc_1...
I checked out the link, and the Sample 2 file does not represent any wave and is not audible, so the article contradicts itself.
https://en.wikipedia.org/wiki/Dynamic_range#Human_perception
If all research is wrong, I'm gonna start drinking vinegar and building perpetual motion machines :P
"Research".. sponsored by corporations, and peer-checked by scientific voting rings. A bunch of incrowd elitists who like to use jargon. Science and politics these days are pretty similar
It’s trivially refutable by placing a 60 Hz strobe (e.g. old fluorescent light or even some aftermarket headlights) at the corner of your vision.
Also, for interactive systems, 16 ms is a large chunk of our reaction time. You need close to 1 ms response times (1000 fps) to approximate pen and paper.
A simple google on 60 fps will still show these “scientists” who claim that we can perceive anything higher than 30-60 fps.
“Science” does NOT equal truth.
As far as perceiving images goes, there's a study at [1] which shows people can reliably identify images shown on screen for 13ms (75hz, the refresh rate of the monitor they were using). That is, subjects were shown a sequence of 6-12 distinct images 13ms apart and were still able to reliably able to identify a particular image in that sequence. What's noteworthy is this study is commonly cited for the claims that humans can only perceive 30-60fps, despite the study addressing a completely separate issue to perception of framerates, and is a massive improvement over previous studies which show figures as high as 80-100ms, which seems like a believable figure if they were using a similar or worse methodology. I can easily see this and similar studies being the source of the claims that people can only process 10-13 images a second, or perceive 30-60 fps, if science 'journalists' are lazily plugging something like 1000/80 into a calculator without having read the study.
There's also the old claim [2] from at least 2001 that the USAF studied fighter pilots and found that they can identify planes shown on screen for 4.5ms, 1/220th of a second, 1/225th of a second, or various other figures, but I can't find the source for this and I'm sure it's more of an urban legend that circulated gaming forums in the early 2000s than anything. If it was an actual study I'm almost certain perception of vision played a role in this, something the study at [1] avoids entirely.
[0] 'Humans perceive flicker artifacts at 500 Hz' https://pubmed.ncbi.nlm.nih.gov/25644611/
[1] 'Detecting meaning in RSVP at 13 ms per picture' https://dspace.mit.edu/bitstream/handle/1721.1/107157/13414_...
https://journals.sagepub.com/doi/10.1177/1477153512436367
Note that 2kHz flicker requires 4000fps to be displayed as video.
The audio, on the other hand, that reaches your ears comes from an analog source, even if it ends up digital in between. There aren't some resolution arguments to be made here, all that matters is that the output device can accurately reproduce the proper analog signal. Which has been proven time and time again, and that any simplification of said signal is imperceptible to anything but the most finely tuned listening devices (or maybe some special "golden ears" that the vast majority of audiophiles don't belong to).
> 120dB is greater than the difference between a deserted 'soundproof' room and a sound loud enough to cause hearing damage in seconds.
> 16 bits is enough to store all we can hear, and will be enough forever.
Correct me if I'm wrong, but isn't 16 bit = 120db about the levels of gradations of sound? Even a 4 bit = 16 levels of sound pressure/SPL could go from 20db, 20+12.5=32.5db, 32.5+12.5db and so on until 120db.
Then, the important question is what's the minimum SPL difference perceptable (at a given spl level). That may well not be 1db.
- Sonos
- Airpods
- Beats
For that price range, Hifiman produces pretty good planar headphones. The edition XS sounds really good.
For a contrived example, imagine an A/B test where you have to tell me which image has more red. Image 1 is a dark red panel on the left and a fully bright white panel on the right. Most people would say the left is more red, but in my fictional test it is actually the white panel because (100, 0, 0) has less red than (255, 255, 255).
If you use ABX, people know exactly what they are supposed to be matching.
In the acoustically prepped monitoring booth of his recording studio, a friend of mine tried to give me an ABX test of 24bit 96 kHz recording and its 16bit 44.1 kHz rendering that was supposedly done right. I heard the difference and easily picked the high-rate one that sounded more life-like. With my best effort, I described it as having a clearer high frequency spectrum, while the other sounded muffled in comparison.
I am left wondering if the 44.1 kHz file wasn't actually rendered correctly with dithering, or if my friend failed to actually get his studio equipment to play it back correctly. I.e. was some overly aggressive low-pass filter done during the conversion or during the playback.
You're better off spending your money on a bog standard DAC/AMP (feel free to opt for tube even, if you insist) combo running through a pair of decent headphones off of 320kbps MP3/AAC (or FLAC, if you insist) source. Even, if we took your subjective insistance that this specialty equipment improved your experience by .00001%, it's probably not worth the 500-1500% increase in expense.
As to your specific example, I can guarantee you that your Bluetooth codec (LDAC or not) introduced far more sound artifacts than the difference between 16 and 24-bit sound.
Also, what is xm1000w3s? I can't find any record of this so I'm guessing maybe it is referring to the WH1000XM3 headphones? Given ldac is also mentioned this seems a reasonable guess as it's a bluetooth model. If that's the case I wouldn't call it "good listening equipment", the default frequency response curve of the wh1000xm3 is incredibly bad, it's barely worth listening to classical music on without using AutoEq[1] or something equivalent (I have a pair and it's much worse than my old Ath M50s which were like half the price). The bass heavy curve of the headphones is far more noticeable than any difference between 16/24 bit audio would ever make.
[0] https://people.xiph.org/~xiphmont/demo/neil-young.html#toc_d...
I'm not an expert, but one claim I saw somewhere is that a higher bit width and sample rate is good for people who are mixing and doing audio processing, even where the final result might get downsampled to 44100 hz and 16 bits per sample at the last stage.
That is all intermediate formats and doesn’t really say anything about what is best for consumers like the standard mastered cd quality at 16 bit 44.1 khz.
Bandcamp is a cool market because I can download wavs from albums to store on my phone. You can see what people use as masters and its all over the place. There are many 96khz masters around and 24 bit depth is popular.
I have a usb audio IO that supports 192KHz across 8xin+out. Those file’s just clog up hard drives so I figure 96 is good enough for bat music.
I’ll run youtube rips of dj sets through some light hardware compressors and preamps and it sounds great. You cannot have specs determine quality.
If proper precautions are not taken during the recording/mixing/mastering phases aliasing artifacts can be heard in the recording. This may account for the differences that some people hear when judging whether there are differences between the two. Higher sample rate files are more permissive of aliasing and exhibit less perceptible artifacts. So you're less likely to hear it at a higher sample rate.
The artifacts of aliasing manifest as inharmonic distortion that starts at the top octaves and then folds back into lower frequencies as the effect is intensified. This can be easily perceived by most listeners if it is pointed out to them. It is not a pleasant effect like first-order or second-order distortion. It does not compliment the record at all.
That said, if proper precautions are taken to mitigate latency artifacts during the record-making process then a listener shouldn't perceive any difference between a 44k and an 88k record. The best case scenario is often a record that's recorded, mixed, and mastered, at high sample rates, even if it's ultimately be down-sampled to CD quality (44 kHz).
So the only situation where 44k and 88k can sound wrong is if... the 44k file is different and wrong?
I give that long preamble to say once a record is done and mastered, having > 16/44.1khz is wasted bandwidth.
If you get silence, they're perfectly identical.
In audio recording, sampling at 88k would be like generating MSAA x2 image, so it can be displayed with higer fidelity, despite the outgut resolution being in lower 44k sampling rate.
The mastering discussed higher up in the thread is going on ahead of time at the studio, not on your playback system. The whole mastering pipeline starts with some initial capture resolution from microphones/cameras. The studio processes these original raw captures into a combined form and prepares the distribution format, i.e. a planned audio/video stream resolution. The studio can use different resolutions during capture, processing, and final distribution.
Generally speaking, the highest rates would be easiest to work with and avoid perceptible artifacts. But practical tradeoffs are made to save cost whether in processing, transfer, or storage.
The "frequency response" spec listed on speakers will tell you what range they are designed to reproduce. Typically, it's approximately 20Hz to 20,000Hz to match human hearing, perhaps with a higher floor if the speaker is designed to be paired with a subwoofer. This range is usually a deliberate (and sensible) limitation imposed by the electronics, not necessarily the materials or the magnet design, etc.
Some speaker manufacturers will list abnormally low or high range numbers on the spec sheet in an attempt to attract customers who mistakenly believe a wider range means the speaker is better. But even those speakers have a steep roll off curve at the extreme ends of the range, so it barely makes any difference.
Example: https://www.soundandrecording.de/app/uploads/2020/10/8361-FR... from https://www.soundandrecording.de/equipment/genelec-8361a-3-w...
But 44.1kHz worked better in the lead up to CDs (works for modulating onto video tape in PAL and monochrome NTSC), so it won until DVD audio brought 48kHz to the masses.
Then I learned that DVD was 48 kHz and not 44.1 kHz so the conversion program I used didn't account for this. I went back and used a polyphase filter to adjust the sampling rate. It sounded normal again but there were some audio glitches at various times.
I went back about 10 years ago and ripped it again and converted it to FLAC which supports multiple sampling rates like 44.1, 48, 96, whatever and now everything sounds good.
Before that, 30 was totally fine for the masses. In fact it was preferable. Cinema is the last big holdout and, apparently, it's going to take at least another decade before even mere 48 is standard. As someone who has been riding the 120+ fos for over two decades, going to the movies is awful, especially action scenes and panning.
I'm not sure what you're getting at with the 120 fps comment, because that is obviously not the frame rate of the finished product, so it's not the same conversation.
The 120 fps was regarding games. While movies are passive, they could still benefit immensely by doubling to 48. Not every scene in a movie is people talking and this is where 24 stops being adequate. Even YouTube has had support for 60 FPS videos for years.
I know it's not a win for the movie industry. They ought to hate it, especially the artsy types.
Don't know if cinema will ever drop 24fps. The shift to higher frame rates is of questionable benefit as it just makes movies look like TV shows. It seems 24fps is what makes a movie feel like a movie.
As for anything beyond 16 bits amplitude on line level, no, you cannot hear a difference. For such a low-voltage signal the resolution at 16 bits is so fine that it already drowns in all the natural noise and THD in the cables, in the amplifier, in your speakers/headphones etc.
It explains why you’re wrong in easily digestible terms & how a 44kHz sample rate will accurately encode signals right up to the Nyquist limit. The second video is an end to end demo showing the process in action.
Thanks a lot for those videos, they were absolutely excellent. For anyone wondering they're presented by "Monty", the guy behind the ogg container and vorbis codec. I probably understood 10% of what he said but that's still a lot.
This is mathematically false. A 6kHz or 8kHz or 10kHz or 20kHz signal absolutely can be perfectly preserved with a 44.1kHz sample rate. Not just kind of preserved, but perfectly preserved.
8:43 in the second video, he goes in to showing what increasing the bit depth gets you.
There's also the issue of the input signal not being band-limited which is necessarily true for real world signals given that you sample for a finite duration.
Maybe you wanted to say square wave?
This just builds on the xiph video someone else linked but essentially
- sine waves are fine as long as you have points for rising and falling edge (nyquist, 44k guarantees 22k sine wave reproduction)
- bit depth only really affects noise floor, so it depends on your audios dynamic range
In other words, the ringing is an artifact of low-pass-filtering on the square wave, not the sampling process itself. A purely analog system with a 22kHz low-pass filter would have the same "ringing", and your ears couldn't tell either way because they can't detect information above 15-20 kHz. It's not even possible to have a "perfect" square wave in the real world because that would require the speaker cone, air molecules, and your eardrum to teleport instantly from one spot to another -- so you will always have some low-pass filtering, and therefore ringing, on any real-world approximation of a square wave.
When you say
> Before sampling, the square wave is filtered to remove all of the overtones above 22kHz, which are not audible (as that's above the upper limit of human hearing) and would cause aliasing issues in the sampled signal.
Do you mean that the human ear "hears" (for lack of a better word) in sine waves because of fourier transforms? So, a pure squarewave signal (cut off above 20khz due to human ear limits) sounds identical to an "imperfect"/wavey sq.wave made from harmonics lacking higher frequency parts?
Your ears don't experience pure square waves. They can't. They'll get the approximation of a square wave as close as they can experience them, but your ear drum doesn't immediately warp from one point to another. It gradually moves. The fluid in your ear has its own springiness. The hairs which do the final detection don't just instantly move either, they're being vibrated by the motion of the fluid in your ear.
And yeah, technically your eardrums can be moved by waves higher than 20kHz. But its not just the motion of your ear drums that give you hearing, its lots of tiny hairs in your inner ear that resonate at different frequencies that gives you the detection of certain audio frequencies that are present. Normal humans (read: practically everyone) tend to only have the equipment to accurately sense up to ~20kHz pressure waves with our ears as sound. As you age the areas which detect higher frequency sounds get less sensitive first, so you start losing the ability to hear those higher frequency sounds first.
It does seem like you're missing a bit of knowledge about signal theory though. That would really help you understand what I mean when I say a real square wave has infinite bandwidth. A very rough and basic idea to help here is that a wave can be thought of as a sum of fundamental sine waves. So a square wave is essentially the sum of all the component fundamentals, each fundamental gets sharper and sharper edges of the square. But the only way for the sine wave to have a truly vertical edge is to be infinite frequency, right? And what are the edges of a square wave? Vertical lines. So to keep adding these fundamentals together to achieve a square output, you'd need to add an infinite sum of sines together to make an actually perfect square. Monty touches on this in that video a bit, but it goes pretty quick.
I understand how a squarewave needs infinite bandwidth for decomposing into sine waves, but a wave like the semicircle one has a vertical tangent and would not need an infinite bandwidth. Btw you're right, I've never had any signal theory classes (studied mech engg).
FWIW, the graph at the top of that article (as mentioned in the comments) does not have vertical tangents. They're not true semicircles. Most of the answers given to make actual semicircles become non-continuous signals (y >= 0 if ... answer) or infinite sum. The one which doesn't just says they look like semicircles. I don't really have the time (or immediate knowledge, I'm admittedly bad at math) to dig all the math but my gut instinct suggests those aren't true semicircles and don't truly have vertical tangents.
And this is kind of how a signal generator can get away making square waves and sawtooths and what not; internally its not always truly "discrete" signals its just quickly flipping a switch from one thing to another. It flips from the high state to the low state fast enough that for your 50MHz oscope it looks pretty continuous, and its probably designed to try and draw a connected line.
This all kind of makes some sense when you get into how we actually make electrical signals. We're normally just modulating the electric "vibrations" of some crystal or accelerating some magnet through a loop back and fourth. These things move in continuous waves. As mentioned, things don't have infinite acceleration they take time to shift between states. So you're never going to get something that goes from high to low in zero time. And most home owners will tell you there's no such thing as a right angle.
I'd like to mention though, you're imaging the wave as being "broken up" into sine waves, which IMO isn't quite the right way to be looking at it. Its not "breaking up" the signal into sine waves, the signals were always sine waves. Remember my comment, truly square waves don't really exist. We can have things that kinda look square-ish when you squint your eyes, but they're not really square waves. A truly square wave in reality requires infinite acceleration or its not continuous. Fill a bathtub and try to make square waves. Its just not going to happen.
I know I'm not fully answering your question, I don't fully know the answer myself. So far in my stumbling around the closest thing I can answer is because that's just how nature and waves are, at least as much as our monkey brains can reason about them. Think about a string in a guitar or a wave in the water or throwing a ball. They all moves in ways which can be described by sine waves. Its just like the nature of how things accelerate and move, changing between states. Why do we see the golden ratio in so many places? Why are circles so special? Good luck on digging for more truth.
The Nyquist-Shannon sampling theorem says that you can perfectly reconstruct any sampled signal as long as the input signal was bandlimited to 1/2 the sampling rate. Using a low-pass filter to bandlimit your 22 kHz triangle wave will remove all the (inaudible) overtones, leaving you with a single 22 kHz sine wave as input to your ADC. The reconstruction filter on the output DAC will then output a perfect 22 kHz sine wave, with the correct amplitude too!
We played a range of snippets of music - rock, classical, electronic, pop - at various qualities over what was quite possibly the best sound system in the world.
The audience was a significant number of record label executives, distribution execs and general audio/music industry experts.
We played pairs of the same snippet and asked people to tell us which was higher or lower quality.
One person got them all correct. Turned out he’d mastered one of the early tracks we played so had a good reference and then used that as a baseline for the others.
Everyone else it was completely scattershot.
It wasn’t a controlled experiment but it was definitely interesting.
This test didn't measure what you probably wanted it to measure.
The original CD version might have some high frequency stuff that's just on the edge of perception that you don't really know is there, but you can just sense a bit of discomfort when listening to it. After going through the MP3 process and that high frequency is removed because it contributes the least in reconstructing that signal, the resultant decompressed signal might sound "better" even though it's not the original, because you get a high quality reproduction but the thing that led to a slight discomfort when listening has now gone.
In this case, the sound engineer got them all right because he could tell the difference and knew what it was supposed to sound like. The rest of the people maybe could tell the difference or maybe couldn't (which was the claimed result of the test), but in fact, even if they could tell the difference, they had no idea which one was the uncompressed one and voted on which they thought sounded best.
As another comment has noted, it'd be a much better test if there was a "they sound the same" option as well as asking which one sounds best.
Playing A & B samples and asking which one is better/original requires much more from the listener that just hearing a difference between the two. It is possible to hear the difference, but not know which is which as that requires additional knowledge.
To avoid this issue you could:
Play (in random order) original twice and processed once and asking which one was different/processed.
Or play two sequences (in random order), [original, original] and [original, processed] and ask, if processed was in the first or second sequence.
Second option might focus better on short-term memory, because it has shorter sequences (2 samples vs 3 samples per sequence).
This would produce a better measurement of whether the difference is audible or not.
Like I say, it wasn’t a rigorously scientific experiment but it was in the context of a conference about evolution of audio standards and what that meant for audio delivery from labels/distributors to DSPs.
And all I can say what I really, really hear the difference.
Because iTunes version is mastered to sound good in iPods so it has a quite noticeable bass boost all over the album, which is extremely noticeable on my 2.0 acoustic which itself has a good bass boost, so this version sounds quite muffled compared to the FLAC version from the CD (CP32-5043).
But on the go, with my CX300-II / CX3.00 there is no noticeable difference.
If you mean 'better used' then yes, if you mean 'vinyl has a greater dynamic range' then...
I do not mean that vinyl is capable of greater dynamic range then digital. Of course not.
skeptical though i may be, i'm definitely not here to say that "audiophiles" aren't charlatans or anything like that, for the record. and while i don't totally understand the setup you describe in the sense that i don't get why insider knowledge on one track would tip all the rest of them (were the HQ tracks played either all first or all second, or something?), wouldn't one person's ability to completely discriminate between the two encodings seem to be very strong evidence that it is possible to tell the encodings apart? the kinds of differences between master recording and 16bit 44.1 kHz are exactly the kinds of things that would give away which encoding is higher quality, no?
i feel like there is this moving target thing that goes on sometimes, where the strong argument made loudly is "no human can tell the difference", and then the tests are more like "most people can't tell the difference between things that they have no reason to be attuned to well enough to have any chance of picking up these subtle differences.
forgive me if i misrepresent, or come across like some sort of audio quality chauvinist-- i ask all of this in earnest, and without having a strong opinion one way or the other.
Here you go:
https://web.archive.org/web/20080322114622/https://www.stere...
https://archimago.blogspot.com/search?updated-max=2013-02-24...
Starts about halfway down the webpage.
Computer graphics is pretty good, but how does it compare to walking out into a bright sunny day.
Audiowise, I wonder how listening to live music, then listening to something that went through capture and playback end-to-end.
I'll bet there are differences and I wonder where the "bottlenecks" are.
Do you think it matters if I play the song on my $10 cheapo earbuds or on $60,000 Sennheiser HE-1 Summit headphones?
In the early days of digital audio, and before oversampling was possible, the anti aliasing filters were analog circuitry.
It's very difficult to cheaply implement a 20kHz brick wall filter with a 2kHz sideband.
Doing it in 4kHz yielded better results at the cost of slightly faster ADC designs.
I believe this is why 48kHz designs got the foothold in professional audio circles. The analog parts of those designs were WAY better sounding.
Once oversampling became common and affordable, the anti aliasing filters where implemented much more easily in the digital domain.
However higher bit rates and sample rates are needed for multi track recordings so that during the mixing stage and mastering, the fidelity is preserved when _math_ causes rounding errors and what not. Unless, you are using nondestructive editing.
As for listening to the final product, i.e. store bought CDs and their equivalent MP3 and AAC rips... I can often hear the difference in specific recordings, no matter the bit rate, because the perceptual encoding schemes often butcher certain recordings.
For example, on RUSH's Red Barchetta from Moving Pictures, there is a synth intro that slowly vamps in volume. Every MP3 encoder (that is normally worth it's salt) I've ever tried encoding that with outputs a garbled, distorted, electronic sounding distortion. It clears up immediately once it reaches full volume, but during the crescendo it falls on it's face.
The question really isn't 'can you tell', it's 'does it matter', and, well, most of the time, no, it does not, even for lower bitrate mp3's.
There are many people, of course, that don't like the idea of lower quality audio, and they can tell at least sometimes, so they 'dislike' mp3 in general. That's all well and good until they start saying silly things like 'mp3's sound bad', which is not true in any sense.
I find it hard to believe an entire industry exists with high-end audio equipment ($100k+ on speakers/receivers/room treatment) just to play 320kbps MP3s?
While I agree with the premise this also depends massively on the equipment used to reproduce the sound. If you have a good amp and large speakers you are more likely to notice than your cheap headphones or thru your crappy laptop speakers.
On a plane, I’m glad I can pull down a ton of albums on Spotify and listen to them without access to internet connectivity.
For Eno, Steely Dan, Roxy Music and Thomas Dolby produced albums, I’m glad for lossless.
For sitting at home on a Friday evening with a bourbon in my hand and the lights down low, nothing beats my vinyl player.
I could get into the cars side of it too but that’s more a “me” thing than a “relevant to this conversation” thing.
> participants were asked to listen to both versions as many times as needed and to choose the version they preferred in a double blind A/B comparison task
I have trouble finding high quality studies comparing 16-bit to 24-bit audio. This one one is kind of interesting:
https://www.researchgate.net/publication/338989993_Study_on_...
2. MP3 has improved a lot over its lifetime. LAME was already used for default by year 2000. When people say MP3 was good enough, they refer to MP3 encoded with LAME. ( Rant: When we people learn the codec, encoder and the encoded results are different things? 2023 and I see this mistakes everywhere still )
3. Even iTunes AAC has seen lots improvement since 2008. Especially in the 256Kbps+ Range.
4. And when AAC is mentioned. That is AAC-LC ( Or AAC Main Profile which isn't all that different ). AAC-LC ( Low Complexity ) has been declared as Patent free by RedHat. There is no reason to use MP3 today.
5. The definition of "CD-quality" alike went from MP3 128Kbps to now AAC 256Kbps. And arguably that is true for consumer market. Even Hydrogen audio has repeated these test multiple times.
6. I still prefer the codec MPC, Musepack (https://www.musepack.net). Sorry I just had to write it out. Sadly it never gained any traction.
7. If we have to be picky about frequency range, may be CD itself isn't good enough and we could use SACD?
8. Lossless is making a come back. Storage and Bandwidth cost continues to fall. ( Arguably not true for NAND, but let's ignore that part for now )
9. It is ironic when Lossless could gain and be used mainstream, Wireless earphones are replacing traditional earphones. Meaning your music will be re-encoded before it is sent to your earphone. And No. Most Android or iPhone dont have AAC pass through. i.e Your AAC encoded files will still be re-encoded before sending it your bluetooth earphone.
It's certainly the convention on HN to put the year in the title for older articles, but it's not one of the guidelines (https://news.ycombinator.com/newsguidelines.html).
(minor point but I can't help it)
Back in the early 2000's when I was getting into ripping my collection I didn't have enough space for FLAC so I surveyed the options and Musepack seemed like the obvious lossy codec winner. I still have that collection of .mpc's somewhere.
Why not then? Because there is a ton of science and empirical evidence that humans cannot hear the difference[1]. Good engineering is about meeting the requirements with minimal cost. If the requirement is that it sounds good to humans, and the cost is number of bits to encode (and thus store and transmit) the signal, then modern codecs like Opus are clearly superior to uncompressed and losslessly compressed signals, much less higher sampling rates.
If your goal is something other than good engineering, for example the aesthetic satisfaction that the bits are the same as what the mastering engineer put on the CD, or for some reason caring how clean spectrum plots of artificial signals look, then the arguments may have some merit. But let's be clear on the goals.
You can compress it for listening later, but you can never add information _back into_ the file. Store it in FLAC for archival purposes.
An equivalent would be archiving works of visual art in JPEG and not something lossless.
[1]: https://www.sweetwater.com/insync/hear-effects-dithering/
Only for some people -- the upper limit of human hearing varies between 15-20kHz, depending on the person and their age. For many children and younger adults (myself included), CRT coil whine is well within our audible range, as an incredibly annoying high-pitched squeal.
This comes up in speedrunning communities sometimes -- many runners prefer to play on CRTs due to their fast response time, and streamers who use CRTs need to remember to set up a notch filter on their microphone, or else their stream may be borderline unwatchable for younger viewers and the streamer might not even realize it.
Because it can fully reproduce everything the human ear can hear. Higher bitrates are only useful for production or archival.
But not because it sounds better. Simply because 48kHz is what your computer and phone natively clocks its audio codecs at. It's done because that's generally an integer fraction, but the why doesn't matter as much as the fact itself; PC "HD Audio" and phone codecs are 48kHz.
Yes you can resample, and yes you can resample without it being audibly noticeable. But it's an extra step where you're at the mercy of whoever implements it to do it right. Doing it wrong may also include noticeable delay, breaking e.g. A-V / lipsync.
Can you build HiFi systems that support 44.1kHz and maybe dynamically switch their clock source as needed? Sure. But what's the easiest way to build a HiFi system these days? You just stick an off-the-shelf embedded device in it, which likely uses standard PC/phone tech…
So just ship 48kHz.
(Similar argument for video recording in 50Hz countries btw - unless you are recording for TV/broadcast, you should always shoot at 60 / 59.94fps. Because that's what PC and phone screens run at…)
Even if it's the only stream and you could switch the codec to 44.1kHz mode, what do you do if the OS wants to play a random notification sound? Switching between 44.1kHz and 48kHz is not going to be hitless on a significant number of HW (not all, but most I'd guess), so whoever's writing your OS mixer code would reasonably make a call to always mix at 48kHz…
(Yes this argument primarily applies to PCs and phones, hopefully on a HiFi system that just happens to use COTS embedded devices they'd write some code to switch the rate…)
The question is if humans can hear the difference in the lossy waveform and if that harms the listening experience.
I recommend that, for serious listening (for some weird definition of "serious"), go to a music concert. PCM is also a lossy compression due to the quantization step, albeit its effect is much less pronounced for so many reasons that no one even thinks it as a "compression" method. If you can tolerate PCM, you should be also able to accept some good enough lossy codecs---I don't know if that includes MP3 or AAC or Vorbis or Opus or whatever, though.
But for other styles, I don't enjoy concerts for audio quality.
It's usually way too loud, so you have to wear earplugs. I've heard some made for this don't skew audio too much, but they are still a filter.
And then you have to like the balance that's chosen by the audio engineers and they are often not ideal. The voices can sometimes be not loud enough to the point you don't hear the words well, the bass too loud. Frequencies don't all travel the same way, so if you are too far away some things are missing or distorted, etc.
And then there's the noises from other people, the claps, the screams, etc.
And the audio still possibly went through some kind of non-analog equipment.
Not saying that feeling the bass in your whole body and feeling the communicative / excited atmosphere from the crowd can't be enjoyable but for audio quality, I'd rather listen to music in a calm room with some good equipment, at a volume level comfortable to me, when audio engineering didn't have to be live and could be (even) more carefully managed.
> If you can tolerate PCM
Are there people who can't tolerate it? It must not be very convenient.
(Huge caveat to this comment: I listen to music most of my awaken hours, but I'm not an audiophile. I never carefully listen to music, it's usually in the background.)
20 * log10(1.0 / 2**16) == -96db
Much like sampling rate, it produces a range that's most likely outside of the ability for any human to appreciably detect. It's also a constant effect, whereas codecs actually analyze the audio to determine which components of the frequency spectrum it can eliminate.I don't think it's reasonable to compare PCM and lossy codecs this way.
Wrong, especially today. Modern ADCs use oversampling to push quantization noise into the inaudible range, and then filter it out before decimation to standard PCM.
Because the end result is standard PCM, the quantization can be only worse, not better.
Oversampling ADCs push a much greater quantization noise into the inaudible range, and then, by low-pass filtering, reduce the quantization noise to the level of standard PCM.
Oversampling ADCs are not better, they are much cheaper, because 16-bit or 24-bit quantizers with enough speed and accuracy are extremely expensive.
Oversampling, i.e. sigma-delta modulation, in both ADCs and DACs, allows the use of much cheaper quantizers with low resolution, of only a few bits, or even of only 1 bit, and of much cheaper filters, which do not have to be very steep, without degrading too much the quality in comparison with real PCM conversion done at the Nyquist frequency.
No, it's a lossy encoding step. Losing at least some of the information of a performance when you record it is unavoidable. For PCM to be a lossless, when you play it back it would have to transport you back in time to when the performance was recorded and you should be able to touch the performer. You're being silly.
The original article claimed that "of necessity [lossy codecs] eliminate some of the musical information". I don't know how to quantify the musical information, but given another statement that "[l]ess bits always equals less music", it seems to be more or less same to the information-theoretic complexity. But as you have correctly guessed, there are a lot of places where the information can be lost, everything from performer's skills, musical instruments, ADC/DAC processes, and up to speakers. So "information not lost" is not a good argument for bashing lossy codecs, because you have lost so many informations already, just that you haven't noticed yet.
Also I claim the lossless compression exists even after this massive reduction of potential information because you can prove that a specific step is indeed bijective. In fact, even the most of lossy codecs are lossless. They are designed to lose some information at the very specific point so that they can be analyzed. "Lossless JPEG (re)compression" wouldn't make sense unless you realize that those lossless steps can be done more efficiently. I have seen enough people who assumed that this is impossible though...
A lot of rock music lives from the imperfection of audio equipment, people spend a considerable amount of time replicating the behavior of vacuum tubes. Even techo producers like Robert Babicz record to analogue tape machine to enhance the final result.
"Lossless compression is benign in its effect on the music. It is akin to LHA or WinZip computer data crunchers in packing the data more efficiently on the disk, but the data you read out are the same as went in."
...but then recommend uncompressed over lossless compression for "serious listening":
"We recommend that, for serious listening, our readers use uncompressed audio file formats, such as WAV or AIF—or, if file size is an issue because of limited hard-drive space, use a lossless format such as FLAC or ALC."
I suppose decoding speed could matter in some situations, but they said "for serious listening", not "if your system is so slow that it fails to decode the file in real time".
Even decoding speed is doubtful, a 486-100 MHz can decode 44.1/16 FLAC in real time with CPU to spare.
But I doubt these guys are using a Pentium 1 machine to play their audio files so idk. The low end smartphone I had in 2013 could easily play FLAC files, at least in the real time uncompressing and decoding part of the equation. Now if the built in DAC and amplifier could take advantage of that extra data is another thing.
Today's standard isn't "CD Quality" anymore. There is literally no audible difference between MP3 320kbps, which covers the complete range of human hearing up to 22kHz and FLAC which covers all the way to 192kHz, which is lossless. At this point digital audio has surpassed what the human ear is capable of hearing, and any advancements to this is superfluous as far as music is concerned.
The only advantage to raw or lossless formats for music is archiving, as FLAC can be converted into other formats without incurring additional quality loss. For listening, it is now more important to have good equipment rather than a lossless format, and for streaming it is generally preferable to keep bandwidth requirements down.
The only reason I can imagine to continue expanding the capabilities of lossless audio is for scientific purposes and machine learning where the limits of human sensory perception isn't a limiting factor.
Communications is the main reason. There's only so much bandwidth in 44.1KHz.
Although MP3 does have some fundamental limitations that cannot be fixed no matter how much bitrate you throw at them (referring especially to the "Inoptimal window sizes" from https://web.archive.org/web/20120222124415/http://www.mp3-te...). They're not dramatic issues, to be fair, but as long as you go for lossly compression, using a somewhat more modern codec like AAC or Opus would be preferable, unless you absolutely need the maximum compatibility afforded by MP3 (though these days at least AAC support should be pretty widespread, too, plus the patents on regular ["low complexity"] AAC have expired as well).
The article's byline has 2008.
A 2023 update could be interesting comparing the streaming providers' choices, and persistence of choices, now that monthly subscriptions, rather than actually owning anything, are so dominant.
The only reason to have "Hi-Res Lossless" is if you're going to do something besides listening with it... and you can't with Apple's streaming.
If we do ABX test you will find out that people can't even make the difference between the original artist and a cover artist let alone lossy vs lossless. Should we just use cover artists at concerts?
The brain adapts quickly to lower quality be it visual, audio, olfaction or gustatory. Does it mean we should ingest the most we can tolerate because we get used to it so we can run on more efficient/cheap resources/content ?
And you would be surprised by the number of radios streaming at 96kbps.
You can absolutely hear the difference between a bad MP3 and the original. I used to amuse myself and friends by quite reliably identifying the difference, blinded, using a rather bad pair of speakers.
Actual CD audio can also work quite differently than any encoding, as at least older CD drives had an entirely separate analog output cable that connected to the sound card and bypassed the ATAPI link entirely. Levels wouldn’t even be matched.
That all said, these days encoders are much better, and there’s no excuse not to go for 320kbps (assuming you have to use MP3).
What I find more interesting is that there was a period where some people who grew up listening to MP3s preferred the artefacting they introduced vs lossless. In much the same way how vinyl enthusiasts like the colouring of the sound that medium introduces. Which just goes to show that as much of this is down to psychology as it is technology.
I mean, the famous mp3 pre-echo was so common in early 90's that I think part of the listeners would prefer listening to it than to a cleaner sound. It is possible that mp3 influenced how music is composed, mastered and mixed.
That being said and adding the fact that people are willing to listen to music using cheap auricular phones in the noisy environment of their cars and recompressed using Bluetooth, I'd say that the 128kbps mp3 is still a very hard to beat format.
People have strong relations with their musics that they are used to while growing up. Nostalgia are powerful memories and they don't want their music unsullied from something that they grew up with.
[1] https://www.soundonsound.com/techniques/what-data-compressio...
[1] https://www.iis.fraunhofer.de/en/ff/amm/consumer-electronics...
Some codecs only work at low bitrates and preserve only narrow bands of frequencies. Some codecs work only at mid bitrates and preserve wider bands. Some codes only work at high bitrates and preserve only the widest bands; you can't get then to drop more frequencies for better savings even if you wanted to. Opus works on all bitrates and gradually and dynamically removes frequency bands as the bitrate drops. Vorbis preserves more or less the same frequencies as Opus at the same bitrate, but loses frequencies a bit faster as the bitrate drops. MP3 drops even faster. AAC works very similarly to Opus, but can't output low bitrate streams.
To compare codec efficiency you would need to do subjective comparisons to see how often each codec achieves transparency (when people can no longer tell if the sound has been compressed or not) at a given bitrate with various types of sounds. This has also been measured, and it's agreed that Opus is basically transparent at 128 kbps. MP3 needs twice as many bits to get the same quality, so Opus is twice as efficient.
My friend, the vertical axis is literally labeled "Quality", and the horizontal axis "Bitrate". The caption is "The figure below illustrates the quality of various codecs as a function of the bitrate." Quality at a range of bitrates is how codec efficiency is measured.
I'd never heard the claim that "Opus is basically transparent at 128 kbps", but I did find https://wiki.hydrogenaud.io/index.php?title=Opus, which agrees with you: "Very close to transparency". But it also notes, "Most modern codecs competitive (AAC-LC, Vorbis, MP3)", which lines up with the chart.
Early Opus vs. MP3 tests were done with LAME, which is awful. This may be why you're under the impression that MP3 needs twice as many bits to get the same quality.
And the labels on that axis make it perfectly clear what they mean by "quality". It's how much of the spectrum they preserve at that bitrate. If "quality" referred to subjective quality there's no reason why the chart should stop at 128 kbps. It stops there because the fullband codecs don't brickwall the signal past that point. Instead they use psychoacoustics to compress it.
>Early Opus vs. MP3 tests were done with LAME, which is awful.
That's funny, because other commenters say LAME is currently the benchmark for MP3 encoders.
Here: https://wiki.hydrogenaud.io/index.php?title=Transparency it states that MP3 is considered artifact-free at 192 kbps, although here: https://www.head-fi.org/threads/when-is-mp3-transparent-an-a... someone did an ABX test and they could still hear differences more than half the times at 256 kbps. If I take the lower number, MP3 is still 50% less efficient than Opus.
Yes, but what about with music?
Yes as several have written, the piece is from 2008 and it doesn't matter any more.
First, once LAME and VBR came about, I've never been able to tell the difference between my 192K MP3 and lossless files, even as a spring-chicken with expensive equipment. Been "good enough" for a very long time.
Second, since storage and bandwidth exploded I've used FLAC exclusively. Why not? But, have found 24/96+ files on the internet occasionally and first thing I downsample them to 16/48khz and do a listening test. I sure as hell can't hear the difference between those. I do leave the last extra 3.9khz... why not? Incredibly cheap and maybe the kids can hear it. Playable on car stereo and more compact, one third the size.
Finally, a big exception. Techies obsess about compression formats, but they don't matter as much as you think at the high-quality end. I've learned the source, i.e. master recording is more important. Example—rip "pristine" FLACs (or WAVs) directly from an iconic 80s CD. Do a listening test. Compare them with a modern remaster encoded with 192K Lame VBR MP3. The MP3 will sound a lot better and preserve the improved high end details. Yes, more noise but you'll struggle to hear it.
(Caveat—this is assuming we're not talking about a shitty 2010-era "loudness war" remaster but a quality-oriented remaster.)
Was mildly surprised by this after insisting on FLAC for almost two decades. A bit too early, in hindsight. Storage is so cheap now though, it again doesn't matter. FLAC it is, Opus from online sources.
ALAC/FLAC files are pretty small, there's few downsides to going lossless. To be fair, there arent that many upsides either, but you at least skip one recompression step when sending the audio over BT.
Gigabytes are cheap.
Yanni Rainmaker Flac -> 40 MB Yanni Rainmaker Mp3 -> 3MB
More than a factor 10 for a single song. For 50 songs that would become 2Gb. I love flac as a format but i would never recommend it as a general format for my grandmother.
If you haven't looked in five years (like myself) I recommend doing that. No one needs to suffer on short disk space any longer. Don't know what "grandma" uses but it is unlikely that audio is a significant burden anymore when people routinely shoot HD+ video.
Also if compressing, Opus sounds better and is smaller.
And I get that the 1TB for $100 is cheaper per gig, but if I never even needed those gigs in the first place my overall cost is still cheaper to get the $3 one.
The last one sadly is from 2014: they tested Opus, AAC and Ogg Vorbis at 96 kbps against a classic MP3 128 kbps, and find out which codec produces the best sound quality.
https://listening-test.coresv.net/results.htm
https://listening-test.coresv.net/bytrack/index.htm
Notice that it is almost 10 years old, and that MP3 was encoded at 128kbps.
Old old OLD argument here. Apart from -rare-, well-trained golden ears, very few humans can distinguish 192K or better, well-encoded MP3s without knowing -exactly- what to listen for. Or will need to 99% of the time.
As dozens of studies have shown over several decades. The rest is either marketing or self-deception.
I certainly don't have golden ears; I'm no audiophile, and I'm getting on in years. 44KHz FLAC is easily good enough for me. But I tire of listening to MP3 music, after a few tens of minutes; it seems to lack the presence and immediacy that keeps me interested.
Not really. I'm not proposing a hypothesis that needs testing; I'm just reporting subjective anecdata. I don't need to test it, because even if I'm deluded it costs me 300GB instead of 100GB. Pfft.
"Lying to yourself" is silly talk; that implies that I'm knowingly telling myself a falsehood, which doesn't make sense. At worst, I'm mistaken.
You are aware of a common fact backed up by mountains of empirical data, a strong physical explanation, and fundamentals of information theory, but tell yourself this lie:
>I'm convinced that we can "hear" frequencies well above the reputed 20KHz limit of human hearing
This is a lie. You tell it to yourself. You must see this.
It's all a painfully fruitless effort when you learn that most masters don't even consider the phasing of instrument microphones and none of it is at all a close approximation of what it would be like to be in a room listening to instruments. It's good enough, yeah, but there are much more important and difficult threads to tug than lowering noise in the signal chain.
I think my point is that for people who have a hi end setup and are used to listening to music -- don't come under the classification of "most people".
I agree with you for that for the average listener with $200 headphones or a club DJ, MP3-320k is fine.
EQ-ing and mastering could also be different.
Possible, although I didn’t change any default ones.
If you see such a huge difference across all music the playback software have manipulated the audio.
Assuming you have high quality set on spotify (even in the mobile-streaming setting, if you didn't use wifi).
Feel free to make one or take an existing one like https://abx.funkybits.fr/test/the-eagles-hell-freezes-over-h...
This is why variable bit rate was developed.
A more modern codecs like AAC or Opus, that can e.g. better deal with the cymbals problem you mention.