Does pono's 24-bit audio matter? Hear for yourself.
blog.beatunes.com
blog.beatunes.com
When you apply negative gain to the 16-bit signal, it has less amplitude in absolute terms, but it should have all the resolution of a whole 16-bit sample in that space.
The point is that whether that 16-bit signal is represented with the lower 4 bits of a 16-bit sample, or the lower 12 bits of a 24-bit sample, you can scarcely hear anything at all, let alone a qualitative difference between the signals.
192KHz is not a higher quality format, it is a production format. The purpose is to reduce the cost and finality of anti-aliasing filters when sampling the signal, it is just cheaper to manufacture. There's a good reason for this being the standard in professional audio, they need to buy a LOT of audio interfaces, many studios will have thousands of such inputs.
Thanks to the basic laws governing signals, we know with near certainty that not only is 192KHz overkill, but so is 48KHz, and so is 44.1. Without significant new evidence showing humans hearing signals with frequencies greater than 24KHz, you will not make any convincing argument as to why we should go with any sample rate higher than 48KHz for human listening.
As for 24-bit, it is another production interchange format. It's there so that you don't need to stand around adjusting gain knobs on audio interfaces so that you get decent fidelity but also don't clip. With 24-bit you can just sample your audio once, and assuming it's within a reasonable range, you can adjust the gain in the discrete signal. There is some indication that 16-bit sampling is less than completely ideal. The very best ears in humanity(newborn ears) distinguish about 21 bits in the safe ranges of amplitude, 24-bit may make sense in an audio system for newborn babies.
I feel that in 2014, music should be released in the best format available from the studio, and if you want a crappier version so that you can shove only 100 songs on your iWhatever, that's your choice.
Wastes space? In an age where we stream gigabyte movies? Do we need to remind the author that we're no longer in 1982 and that we have evolved beyond floppy disks?
Clever way of leaving out just how much space it takes up (takes up 6 times the space [0][1]) also as both this author and Xiph concluded you can't really hear the difference. When we stream HD (As you put it "Gigabyte movies") we are getting a clearly superior product. I can easily see the difference between 480/720/1080 whereas I seriously doubt I can hear the difference between 16 and 24bit audio. Yes space is cheaper than it used to be but it's neither free nor unlimited, not to mention the primary way people listen to music nowadays is probably on mobile devices where space is still at a premium or they enjoy their music via a streaming service a la Spotify/Rdio/Google Music where bandwidth is at a premium (And storage factors in here as well due to either being able to store less songs locally or in cache due to larger file size).
[0] http://people.xiph.org/~xiphmont/demo/neil-young.html
[1] http://blog.beatunes.com/2014/04/does-24-bit-audio-matter.ht...
filesize(192kbps + 16bit) = 6*filesize(192kbps + 24bits)
Also the point of both articles is that: You can't hear the difference.Look at MP3 vs WAV (or FLAC), MP3 won out (for consumers) because of it's filesize being so much smaller. Look at the compression differences:
>> Uncompressed audio as stored on an audio-CD has a bit rate of 1,411.2 kbit/s,[note 2] so the bitrates 128, 160 and 192 kbit/s represent compression ratios of approximately 11:1, 9:1 and 7:1 respectively. [0]
At the best quality listed here (192kbps) MP3 is 7 times smaller than it's WAV counterpart. I just don't see people paying for 6x the space for something they can't tell the difference between. If anything history has proven this not to be the case.
192 ÷ 48 = 4
24 ÷ 16 = 1.5
4 × 1.5 = 6
The six times is also on the wrong side of your equation, I guess.But if you need a source, I can cite what you cited:
> Unfortunately, there is no point to distributing music in 24-bit/192kHz format. Its playback fidelity is slightly inferior to 16/44.1 or 16/48, and it takes up 6 times the space.
They were talking about both sampling rate and bit depth.
Oh, and just for the record, I do agree with the general idea that it's useless to ship audio to consumers as 192/24, for the same reason it's unwieldy to publish photos online as 20+ MiB raw files (even if browsers could display them). I was just noting that the size comparison included both sampling rate and bit depth and not just one of them.
filesize(192kHz + 24bit) = 6 * filesize(48kHz + 16bit)
It also only applies directly to raw audio like pcm. Once you encode it (to, say, flac) some of that size increase could disappear depending on the waveform.
I don't know how mp3 is defined for higher dynamic range, but I know frequency range/sampling rate doesn't matter, because a 128kbps mp3 is 128kbps whether the frequency cutoff is 14kHz or 20kHz. Isn't a 192kbps mp3 from a 16/48 master the same size as a 192kbps mp3 from a 24/192 master?
With the references to 192kbps and mistakes, it wasn't clear what parent was talking about.
Betamax didn't lose to VHS because it was worse.
Well, he probably also knows that we now increasingly use SSD disks for our laptops, that are more expensive and sold in lower capacities than HDs on average.
Or that we also listen to music on mobile phones, with like 16 or 32 GB or space (which we also want for apps, photos, videos and other stuff).
Or that we might have "evolved beyond floppy disks" but we have also evolved beyond having 20-100 albums in our collection. In fact digital music collections of 1000 or 10.000 albums are not uncommon.
Now, 1000 albums of 500MB each would make it 500GB for an music lover's collection. That's a non starter in 99% of modern laptops, and impossible on an mobile phone.
I'd rather have it compressed more or in less than 24bits, and fit more stuff to have with me to listen.
(Also keep in mind that music lover is not the same as audiophile).
Don't use space that is literally wasted, because it adds zero information.
It's equivalent to a database which adds a sequence of 32 bits of 1's after every field. It adds nothing for any purpose, yet it uses space. This space is truly wasted.
Same with storing audio we cannot possibly hear.
Honestly as an end user its really good enough. If I was remixing this stuff and processing the audio I would want the non-compressed. (similar to shooting Raw photographs ). But at 256kbps honestly it really is amazingly good. thinking back to cassette tapes and records, its astonishing.
As for the tape hiss on the classical music, the performances are so good I just ignore it. Sometimes I think people get worked up about the quality of the recording and ignore the quality of the performance.
This is the main problem with audiophiles.
Improve your equipment until it makes your music awesome, then just enjoy it.
Then again, it's been proven time and time again that the average person can't even hear the difference between mp3 or FLAC.
Considering you'd also need proper high-end monitor headphones, i believe this is more about marketing, the brand and the pitch than it is about what you actually get or can actually use.
+ it can only hold about 2000 FLAC files even with expansion card.
I always like Alan Parson's quote, audiophiles use your music to listen to their equipment
Of course, few people actually want impressive shifts in volume in their music, because prog is dead and no one wants their headphones to suddenly blow their eardrums out. Gotta compress the shit out of that dubstep for the kiddies! Somewhat ironically, greater bit depth in audio makes a bigger difference for film than most music because of this.
The design is not for everyone, but I think when viewing it through this use case, it has a much more successful design than a flat form-factor like most portable music players
But... they could have just made it flat, and then it fits the use case you laid out AND the pocket one. Like, there is zero advantage to the bulky shape.
It isn't equivalent enough to real music, and the human auditory sensory perception and recall systems just aren't perfect enough to make accurate judgements in those test cases. When testing you end up looking for discrepancies you can identify with words and identifiable momentary observations, things like "I can definitely hear this passage the notes are more muddled together in the compressed file." But you miss so many things that show up as minor feelings of unidentifiable hunches. Of course, you could change the format of the test, but even still, "identifying" is not "listening."
Things you can not identify, or talk about, or remember, or form sentences about, can still impact your musical experience. Our ears and brains are complex, and the range of input that is perceived subconsciously is astounding. We accept this for vision, for taste, for emotions, for memory, but for some reason not for audio. The simple feelings that you can't put your finger on can be important to the experience even if they can't be identified. But if you start to talk about "realism" and "emotion" and "feeling," you are immediately blacklisted by the empirical measurement mafia.
There is so much truth to the fact that scientific measurements and empiricism are important in determining the decisions you make about your audio storage and listening. You shouldn't pay money for things you can prove won't make a difference, and there are so many things out there that fit that bill. But we shouldn't throw out entire possibilities of discussion just because they influence parts of our experience that are subconscious or unidentifiable in an A/B test.
Again, I'm not saying we should disregard scientific evidence, just that listening tests set a limit that might not be accurate. We're finding only things that can be tested, and many parts of the audio experience don't fit the test.
In this case, there are more bits down there, higher resolution at lower levels. It basically means you can turn up the amplification and receive more resolution regardless of amplified volume. But it also means you have more information, and that's worth something.
Significant noise power at ultrasonic frequencies (or even a skewing towards higher audible frequencies) represents a problem for tweeters & amplifiers both in terms of the intermodulation distortion mentioned in Xiph's article and in terms of power dissipation at the business end. IIRC Dell and VLC are having a falling-out because VLC's soft clipping is damaging the speakers in certain Dell laptops.
Dell probably didn't bother putting an analog reconstruction filter on their DACs, assuming they used an average Realtek codec the digital filter in the DAC has a minimum cutoff point around 28khz but could be higher when operating with higher sampling rates. They probably also sent the signal into a filterless class D amp to drive the output. Those two things together add up to a lot of ultrasonic crap in the signal.
VLC allows amplification of the INPUT above the sound that was decoded. This is just like replay gain, broken codecs, badly recorded files or post-amplification and can lead to saturation. ... snip ... At worse, this will reduce the dynamics and saturate a lot, but this is not going to break your hardware.
Except it can - because the saturation skews the power distribution towards higher frequencies which weren't designed for. This is very, very well known - it's the reason why one always chooses an amplifier with higher output rating than the speakers, often by a factor of two. It's counter-intuitive but driving a low-rated amp into saturation can overheat and destroy tweeters. (Guitar amps get away with it by having massive voice coils on a speaker which will only ever need to reproduce up to about 5kHz.)
I barely can tell the difference between high sample rate 8 or 12 bit audio and 16 bit audio. In some circles I hear people talk about that "gritty" 12 bit sound, but its usually not properly dithered, heavily filtered, 20khz or less sample rate and generally destroyed on purpose.
Audio is generally the best arena for snake oil since subjectivity is so high. I cannot believe that Focal Grande Utopia EM loudspeakers(180,000 usd MSRP) exist in quantity, but I know two people who own a set in their (comparatively) modest living room. I have heard them in action, they sound great, but Im not sure that I heard anything in any of the music that I hadnt heard with good headphones or speakers. At what point is it "good enough"? I own plenty of recordings that no amount of hifi audiophilia will help.(many that I recorded myself ;) )
By default, this is not true because sox will use a higher bit-count and perform dithering when converting to a lower bit format (ref http://blog.beatunes.com/2014/04/does-24-bit-audio-matter.ht... ) in order to remove the extra bits while reducing the impact on dynamic range.
In all fairness though, this is the same argument used in the Xiph article ( http://people.xiph.org/~xiphmont/demo/neil-young.html ) in an attempt to justify the use of 16-bit over 24-bit, the prior allowing for a comparable dynamic range with the use of appropriate dithering techniques. I just think it's important to note that the process used with sox is doing more than simple arithmetic bit-shifts/multiplication.
Anyway, nobody here is both qualified and willing at the same time to tell you what you need to know about this topic. The article referred to, written by Monty at Xiph, should give you a very good overview of how this works .
Maybe people who work in professional (not consumer) audio? Who design these systems for a living?
http://www.eirec.com/DPimages/digisqwvtest.jpg This is an example if transient response of different sampling rates.
Sampling theory says that a perfect square wave can be represented at any frequency below Nyquist. That doesn't mean that the codec or the analog electronics are capable of responding instantly at those frequencies, but that has nothing to do with the fact that a 1kHz square wave can be perfectly sampled with a 44.1kHz sample rate. The image is simply incorrect.
Transient response in the real world is generally limited only by the acoustic transducer response of the system, because everything else has the ability to respond much faster than audio rates. With Pono, this means that the earbuds or headphones you use with the player will have a greater affect on the transient response than the electronics inside.
It has been proven over and over via ABX testing that high res formats are completely indistinguishable from a 24b/44.1kHz master. And further studies have proven most audio engineers can't distinguish between lossless audio and 320kbps MP3. That's what the Xiph link elsewhere in the thread shows.
Sure, but when sampling and then processing things, you can consider it "headroom". Just like it's not hard to edit a photo in a few steps that the limitations of 8 bits per pixel begin to show. Though I suspect sampling frequency is a lot more important there, 24bit can't possibly be "too much".
Yes, I understand that this is not a good argument for consumer products to be more expensive and need more storage, just so musicians of the future, or pirate musicians of today, can have a better time. So I think the best plan is to encourage everybody to become hobby musicians, then it would be an easy sell.
But I always wonder how much of the sound we hear by feeling (skin, hair vibrations). Because maybe I can't hear 1Hz, but I can very well feel it.
21-bit would make for an awkward file format. If we can hear 17 bits of range that's enough to justify storing music as 24-bit.
I have very sensitive hearing, I hear things most others cannot. I can detect 192kbps vs 320kbps mp3 encoding. I cannot have any switch-mode AC/DC transformers in my bedroom because the switching noise actually keeps me awake (I charge my phone in the kitchen, most people think I'm mad when I complain about the "noise" :)
Still, I've never been able to detect frequencies above 20Khz or hear the difference between 16 and 24 bit. I think anyone claiming to is full of crap.
24-bit is pretty much useless for playback. While human hearing range may extend past 16-bits of dynamic range, that doesn't mean that full range is musically useful. It is, however, quite useful to record and mix in 24-bit for the headroom.
About the MP3 files, yes, the difference is audible. But of course you need good headphones.
You can't get an 18 kHz square wave out of a system with 44 kHz sampling. You need at least 1 harmonic before it'll even LOOK square, and that requires a frequency response out to 54 kHz, ie. a sampling frequency of 108 kHz. You CERTAINLY won't find one in a reverb tail, even assuming you had a generator for one in the first place (you might JUST get one from a cymbal crash, but I don't think the physics works)
The point being, your source material can't contain an 18 kHz square wave either since it's been through a studio production system with the same antialiasing filters.
Since you know nothing about me but seem to be making assumptions anyway, here's some background. I've worked in broadcast audio; I own studio recordings in 24 bit / 192 kHz (Linn release of Mozart's Requiem, studio master series). I also own studio equipment that can actually play it. Audiophiles are, by and large, cash cows for companies with no scruples.
Seriously, A/B test this, you might be surprised.
Also if you think that's anything like a square wave coming out of a guitar speaker (or that that is even desirable in the most case), I've got a bridge to sell you. And yes, I do play.
Okay here is a picture of what I'm trying to explain. And the author of this picture used a frequency much further inside human hearing range. This is transient response test I guess. My main argument is for the verbatim capture of the input wave. It will make the sound at 10k but it isn't the same wave that went in.
It looks like an 18khz sine wave, possibly with slightly reduced amplitude depending on the anti-aliasing filter rolloff and fc, but not enough to be audible (18khz isn't audible for a large portion of the population anyway).
> How is phase affecte?.
Probably delayed a bit by the anti-aliasing filter.
> Phase of an audio signal reaching the ears helps you perceive distance and position.
Not at 18khz, the wavelength is too short for your ears to notice any realistic group delay. High frequency localization is most done by ILD and effects caused by ear shape.
> What is the quantization error difference between 16 and 24.
Quantization error in a DAC just defines the noise floor. 16 and 24bit DACs are usually within a few dB of each other in dynamic range, it's really not audible.
> And if the author can't hear the difference between 4bit audio and 12bit audio I question what he/she is listening on. The aliasing would be HUGE.
What does aliasing have to do with bit depth?
but that doesn't have anything to do with number of bits? I mean, an audio player system is basically a digital stream sent into a DA converter outputting an analog signal which is then fed into an amplifier where the signal is amplified and then sent into speakers of some sort. No matter if you have an 8bit audio stream, or a 128bit audio stream, the stages afterwards play a much larger role in the hearing damage caused. Also not all DA converters ouput the same level to begin with. Some are -5V to 5V while others (single power supply) can be 0V to 5V, etc etc.
If you set up a 20 bit system (120dB) so that you can hear the lowest bit toggling, 120dB above that will most certainly cause hearing damage if used for very long.
144 dB (24 bit) or more SPL will cause near instantaneous damage and pain.
And that's the point.
When you print a photo in size 9x13cm, does it matter whether you print it in 10,000dpi or 20,000dpi? Hardly, unless you always run around with a microscope.
And it does not matter, whether the photo was originally a good picture or not.
It can matter when you put it in a scanner though. I tried scanning some old photos, they look horrible compared to even my puny 8mp camera. Then I tried scanning a piece of cloth, and was blown away by the detail..
(and I'm inclined to agree)