24-bit/192kHz music downloads and why they make no sense (2012)
people.xiph.org
people.xiph.org
That’s not why I go for High-Res stuff, though.
It’s all about archival, at least for me. With a 24/192 Master in FLAC or ALAC, I can downsample to whatever the destination form factor is. I can transcode to a 320kbps MP3, or a 16/48 WAV stream for a smart speaker, or a 24/96 stream for the theater. The point isn’t that I can hear the difference, it’s the fear that I might lose something irrecoverable by sticking with lower-quality files for bulk storage. Once data has been discarded, it cannot be retrieved, and that influences my preference for storage (and is also why my BD/UHD rips are into MKVs, no re-encoding).
Now that being said, I will absolutely hem and haw and ABX different releases to determine if I opt for the 16/44.1 CD rip of an album from the 80s or the new 202X remaster in 24/192 (spoiler: almost always the former), and I absolutely prefer anything with classic instruments (Jazz, Classical) in higher-quality formats because of a subjective perception of a wider, clearer sound stage, though this is almost certainly a psychological effect from performing in concert bands and orchestras rather than physical or objective in nature.
Like I tell newcommers: if it sounds better enough to you to warrant the purchase price, then that’s all that really matters. Enjoy the hobby.
The takeaway from these sorts of posts, at least in my opinion, should be two-fold:
* Understand the physical limits of human senses and perceptions to help inoculate yourself against outright scams and grifts
* Liberate you from the "tech grind" and allow you to enjoy what you like, how you like it.
Also understand that while there is an upper limit, we are all different within that. I can hear the difference between 128Kbps and FLAC, at least for some content, but not 256Kbps, maybe not 192. For some content (spoken word etc.), 64Kbps, sometimes less, is perfectly acceptable (to me). There was a time I could hear the difference between some encoders, but that was decades ago and anything in active use is pretty damn good (and my ears are not what they used to be) unless you really crank the bitrate down or tweak other options daftly.
You've established this with double bind testing, correct?
Not recently, so it is possible that improvements in encoding methods, and changes in my ears, could mean that I'd get a different result now.
A reasonable definition of transparency for high bitrate compressed audio is "Can the worst files be distinguished by a listener trained in what artifacts sound like". Maybe also add in having to use a high discrimination listening setup, including not running excessively loud (increases masking).
If that's not the test you're doing, it's unsurprising. At moderately high bitrates no one can reliably distinguish them on arbitrary samples: most inputs are easy.
If you test on known-difficult "killer samples" you'll probably easily distinguish them, even without first being shown what to look for, and certainly after.
During the development of Opus I created many 'trained listeners' and selected many killer samples, and I don't recall* ever encountering a tin ear that couldn't be taught to ABX any high rate samples, though some people are obviously much better at it.
I'm not sure I'd recommend it though: learning to identify artifacts has a frequent side effect of making low rate audio like the HE-aac used in SirusXM absolutely intolerable. I'm bothered by it even when I hear cars driving by using it. :)
[*] My memory for such things sucks, so I could be wrong-- but my point that it's not expected remains.
You're right it's just minor details.
I also spent a lot of time ripping my old CDs to FLAC and trying different MP3 and AAC encoder settings to get playback that felt transparent enough to me. I could never tolerate Sirius/XM radio streaming due to the horrid compression I heard with every futile attempt. I still seem to have more sensitive hearing than most people around me, but in my 50s I know it isn't what it once was.
I never had huge budgets, but did strive for hi-fi in my limited ways. I used things like toslink and HDMI to send raw PCM data from Linux to my Yamaha A/V receiver's DACs + amplifier to drive somewhat nice Polk tower speakers. But then COVID-19 happened, and this stuff was packed up to move house.
Nowadays, music playback is streaming with mundane "subwoofer + satellite" PC speakers or MP3 playback with a mini-SD card permanently parked in my car's infotainment system.
As referenced in the article, a common explanation for those audible differences is that the high-resolution version of the album is sourced from a different master.
In fact if you can't hear the difference between 24/192 and 16/44.1 you shouldn't be working in audio. (Doesn't apply to consumers. Does apply to musicians and engineers.)
It's like being colour blind.
And if you don't understand the math behind quantisation, you shouldn't be posting pseudo-scientific videos where you use an oscilloscope and a cheap spectrum analyser - both tools with very limited resolution - to "prove" your point.
16 bit isn't enough for hard, objective reasons. One is that the noise spectrum of quantisation is not simple. Most people assume it's something close to plain white noise, but it really isn't. It's actually a very complex spectrum with some prominent peaks at specific subdivisions of the sample rate. Those frequency peaks are significantly above audibility. 24-bit quantisation shrinks them below audibility.
The other is that most people can hear dither/noise-shaping at 16-bits. That adds a single bit of noise which should - if you're being very literal - be far below the threshold of audibility. But it clearly isn't.
These two facts are related.
The more complex reason is that listening is an active perceptual process. The brain does a huge amount of processing to separate sources and place them in a perceptual field which includes information about perceived object type, distance, and ambience cues. Some of those cues are very quiet, and we don't hear them linearly.
So using sine waves as some kind of perceptual reference for audibility is nonsensical. We hear much more complex signals in an active way, and if there's information missing in the quiet parts - which there is with limited quantisation - then the signal simply isn't accurate.
But it depends what you're sourcing from. If you source 44.1 then you will have a worse recording if you change it to 192. If you source at 48k then you just waste samples. If you’re recording analog inputs at 192k in a crappy adc then you will have a worse outcome than a good adc at 48k (or 44.1k)
Same with bit depth - the adc is far more important.
It's not like all of your samples and virtual instruments are 192khz or even 96k. Many are 48khz or even 44.1k.
I think there are many cases where people never need to go above 44.1khz unless you maybe have saturation on the master bus. I agree that good dithering is important though and think that there hasn't been enough research on that so far.
What you are describing is the result of blunt truncation. If you use the most basic (“uniform” or “rectangular” a.k.a. “RPDF”) dither, the spectrum is in fact flat, as demonstrated by the video you are likely alluding to and calling “pseudoscientific” (https://youtu.be/cIQ9IXSUzuM?t=12m50s). If you sum two uniform dithers together, you get what pretty much everyone uses (“triangular” or “TPDF” dither) which, in addition to decorrelating the mean quantisation error from the signal, also decorrelates the standard deviation, eliminating noise modulation and leaving a correlation only in still higher-order moments like skewness and kurtosis.
You can even try it for yourself with SoX. Find a 24-bit track, quantise it with dither to 16-bit, calculate the difference between both tracks, blow up the difference and take its spectrogram and it will be completely flat. Or listen to the difference (mind the volume) and see if you can make out anything meaningful.
$ sox source.flac -b 16 dithered.flac
$ sox --combine merge source.flac dithered.flac loud-difference.flac remix 1,3i 2,4i norm -1 spectrogram # assumes stereo input
$ open spectrogram.png
$ open loud-difference.flac
And then remember that this difference would normally sit at roughly -93 dB FS, so to hear it in a typical room, you would have to be listening at deafening levels. You claim that it “clearly isn’t” below the threshold of audibility but it’s not clear how you arrived at that conclusion. You then claim that the audibility of that noise floor is somehow related to what you said before about the effects of undithered quantisation, even though those effects stop being relevant the moment you apply any sort of dither.> We hear much more complex signals in an active way, and if there's information missing in the quiet parts - which there is with limited quantisation - then the signal simply isn't accurate.
It’s not missing. You can do a similar test where you “bury” your source material in the 16-bit dither noise floor, blow it up again, and you’ll be able to detect it under the noise.
$ sox source.flac -b 16 quiet.flac gain -100
$ sox quiet.flac loud-again.flac norm -1
$ open loud-again.flacIn this case, it was my brother's own 24/192 recording, down-mixed by him to CD format with the intent that it be transparent. I believe he said his software was supposed to be dithering, but this was ~25 years ago and I can't really confirm the details anymore.
No one can hear the difference between properly mastered high res files. I will happily put money on it.
Small differences in gain are ABX able much more readily than differences in noise at the 16 vs 24 bit level. So if the signal chain gives even a small difference in gain between the samples that's what you'll track. A reasonable conversion path to 16 bits for mastering will also apply dithering and some kind of brickwall limiting (you have to limit after the dither or as part of the dither as dither can change levels!), and this can result in gain changes. The DAC may behave differently or have outright bugs for some configurations too.
This is particularly true wrt reconstruction filters for sample rate differences. And if you were comparing 44.1k and 192k then the physical DAC itself was likely running at a different rate and its _analog_ filters are probably better optimized for one vs the other (this is less true for 48k vs 192k, as the hardware likely runs at the same rate for both). So one answer to this comparison can be "on this particular hardware this rate is better than that rate"-- but that's a implementation property not a property of format choice.
You might think, "okay I'll use a mathematically perfect down and up conversion process and run the DAC in the exact same configuration for all cases". But even then you run into issues like after reconstruction the _inter sample_ peak levels will be higher than the levels of the samples, so you have to handle that and in a way that doesn't produce a gain difference between the two configurations. (probably by running your perfect process and finding the gain level that results in no limiting, then making the gain of the original match).
And then for the high rate vs non-high rate you have to deal with the fact that most amplifiers are not particularly linear (compared to well constructed software at least!) and that any real speaker is very far from linear. This means that the presence or absence of ultrasonics will change the audio in the 0-20khz band.. Before you think "well that could be a reason that high rate is better" observe that if there was some consistently good effect from the ultrasonics you could just bake it into the low rate sample.
> but in my 50s I know
Yeah if you're in your 50's you're absolutely not hearing differences way up above 20khz (especially if you're male), I bet you can't even hear CRT flybacks from 100 yards anymore. :P Most people have no idea how much their high frequency hearing degrades as they age because it plays approximately no role in your life, but it's real, dramatic, and as far as I know happens to everyone.
I don't mean to discount your experience: I don't really doubt that it was real. But answering the general question of the necessity of low vs high rate probably takes a team of experts, armed with test gear and the designs of the HW/SW in question, to vet the test configuration. Testing a _particular_ configuration without the ability to distinguish its implementation quirks from format-fundamentals is much easier and that's what most attempts to test this question are actually testing.
By testing in a recording studio you were doing far better than most such comparisons. Usually people try comparing different files and they're comparing entirely different mastering processes. Files made for the "high res" market will often have much less compression and limiting then files made for commercial radio play / casual listening... and truly do sound obviously much better. Some of my favorite recordings are rips from vinyl. Vinyl is an awful format from the perspective of audio fidelity, but it's also pretty intolerant of excessive compression and limiting because the record will skip if the needle is bouncing off the rails. And more recently I suppose they also avoid over compression there because of the difference in target listener/environment.
This was common knowledge at least as far back as the mid 80s, when every hifi shop and salesguy knew to ensure the bit of gear with the highest profit margin got played an almost imperceptible bit louder than the gear the customer came in to buy during back to back testing.
Point being: it doesn't even require an unscrupulous sales person to get similar results to an unscrupulous sales person! :P
This was supposed to be running the DACs to match the source configuration, not resampling into some common format. I think that is an unavoidable part of the whole end-to-end ABX test concept.
Maybe it would be interesting to up-sample back into 24/192 and play both in that mode. But then people would argue about what type of up-sample to use.
I was in my mid 20s for this test. I understand my high-band hearing was better back then.
Second guessing it by upsampling in front of it seems dubious to me. It might help in some cases where the DAC designers were thinking of different objectives or just didn't do a great job. It might also help with some other issues, like if the dac is timed off the input clock and the input clock sucks and the upsampler retimes the signal.
Of course the upsampler designers could also get it wrong, be aliasing the hell out of the results, and happen to like the sound of the corrupted audio. :P
The effects are all objectively measurable however-- with expensive equipment at least. I think I'd want to set test results with a particular hardware combination before sticking an upsampler in it. OTOH, if there already was one there because it's just some built in feature of some kit I wanted to use otherwise, I wouldn't worry much about it. Particularly if that kit has been reviewed by people with proper test gear and they didn't decide that it was broken.
But I think I lost the ability to hear the flyback not long after I passed twenty. The world turned silent as far as that's concerned (before, you could hear it anywhere and everywhere, in shops, homes, some workplaces..)
The "20kHz" thing is kind of a myth for most people, at least that's what it looked to me after all the testing we did at school. I think it can influence what you hear, somehow, but in any case it's for very young people.
> Most people have no idea how much their high frequency hearing degrades as they age because it plays approximately no role in your life, but it's real, dramatic, and as far as I know happens to everyone.
I agree completely. I recall some discussions a long time ago on RMMGA (Usenet: rec.music.makers.guitar.acoustic) where some distinguished and experienced, but middle-aged guitarists got practically angry when a young guy described the sound of a certain type of newly-introduced strings "harsh" and "like fingernails on a blackboard" when used on a particular guitar.
The difference was, of course, that what the young guy could hear is something which stopped existing at least when you had passed 30.. I was at an age where I too couldn't hear that kind of sound from strings, but it was still not that long ago and I remembered and had noticed the difference, i.e. that I could not hear what I could hear before. For example the huge difference between fresh strings and week-old strings (and that fact has, over the decades, saved me tons of money which I would otherwise have spent on replacing strings all the time..)
You try to hear the brickwall by the muffled, enclosed quality and possibly by the weird pre-ring blurriness of the filter making things sound more vague than they have to be, and you hear the truncation not because it is audible 'distortion' as we know it, but because depth collapses and it sounds like it's coming from the speakers and not being a separate space behind/around the speakers. At no point will it be the most glaringly obvious thing but it'll never be 'distortions' as we imagine them, it's more a 'pod people' lack of personality thing.
Like a much subtler version of listening to AI music :)
I'm quite happy with 24/96 as suitable overkill for anything I might want to hear or do. Neil Young went hard on the proposition that 192 was necessary. Sold the Ponoplayer, I had one but it died on me, battery failed eventually. It really did sound awesome beyond just about any other listening device I've ever heard…
The last couple of generations of converters have gotten a lot better, so 192kHz today is likely to sound cleaner and smoother than it did ten years ago, where there was a good chance the clock was quite jittery.
Personally I don't think it's worth the extra bandwidth for playback, but I can understand why some people might want it.
Generally all of these "debates" come down to people who think math > circuitry. All real designs are imperfect trade-offs. They all have issues, and arguing as if converters are perfect when they never are, and the imperfections can be benched objectively, is... not very scientific.
There is one purely objective benchmark: a true blind test. You can believe if something is different or not, but if nobody's capably of hearing the difference, does it matter?
You can say that and be correct, while also sounding a little more silly than perhaps you'd like.
edit: rather than go even harder, I'm instead going to suggest it's perfectly fine to care about things you don't hear every single time, but still like or dislike :)
My pet example is sand in the lettuce for a salad. If you dislike that particular cronch against your teeth while eating salad, it has a spectacular ability to ruin your enjoyment of your salad, even though you don't perceive it every single time. Digital distortions are like that for some of us, things like wow and flutter and vinyl surface noise are like that for others. People vary. (which is also why not to generalize about what 'people can hear')
There's no reason you can't gather statistics about a representative part of the population. It doesn't make sense to make the entire world pay for better audio because some guy somewhere might be a bit more sensitive and he thinks it ruins everything.
I don't know if such a person exists, but especially in the 'magic' audio territory there's a big amount of bullsh*t going on. There's a reason the James Randi price was never claimed.
There's a huge difference between digital 'distortions' caused by sampling at CD quality and things like tape flutter which most people can actually hear. Even then, some people like the imperfections like vinyl or tape artifacts. Some people even prefer MP3 compressed music.
If you look at a site like audiosciencereview.com and pull up measurements of a DAC or ADC, you can find graphs of the antialiasing filter response. Some are great and some are not.
One could think of 16/44.1 PCM as being a codec that is potentially perfect but requiring some degree of care to encode and decode correctly.
High-dynamic-range material benefits from lots of bits.
But most music today has heavy compressors in the pipeline that kills dynamic range in favor of allowing you to hear almost even whispers, even in traffic or a city with ear pods.
But if you're from the first group, as you said it's more noticable the benefits of having better codecs and bit depth vs heavily compressed top billboard songs where even listening the master track from the studio, falls into diminishing returns.
I used to think the same. But I realized that downsampling hi-res music to 16/44.1 isn't a transparent conversion. So now I prefer the one downsampled to 16/44.1 by an expert in production env. I almost always download 16/44.1 flac files because of this.
Higher rate sampling is just like storing integers to 3 decimal places, or archiving an upscaled DVD.
I recommend you actually read the article. I vaguely recall they did it in video form too.
On a tangent, whenever someone mentions LP sounding warmer or whatever I like to point out that I prefer wax cylinders (a.k.a. phonograph cylinders).
I had a good laugh listening to the sample at https://en.wikipedia.org/wiki/Phonautograph
The sound is very pure indeed.
If I have an option to get a 16bit version of a recording or a high-res version, I choose the highest quality version very time
Same with a physical copy. A limited edition, better quality vinyl LP is more attractive if you are going through the trouble of curating a collection.
I’ve been curating a music library of digital files since before the iPod was released and I will always go for the highest quality version out of principle. I can always downsample it to any thing that makes sense.
But I also have a large multi-terabyte music collection, I follow new music, go to concerts, go to parties, talk about music with my friends in signal group chats.
It's a hobby, and when you get a bit older and start having some savings, if you love music treating yourself with a better system is not that crazy.
Also with HEDD you get a handcrafted device made in Berlin. And if you go with nicer cables, they are very beautifully done and feel great. There is no difference in sound of course. Some people like jewelry, I can get similar enjoyment from beautiful audio equipment and cables.
And so many CDs of course.
It may be simultaneously true that:
A) Humans cannot tell the difference between 44.1kHz/16-bit audio and any higher resolution, and
B) For a particular song, the best commercially available 44.1kHz/16-bit version may not be the best commercially available version
"The quality of the particular mastering can still make a noticeable difference, regardless of the ability for the digital sampling rates to perfectly represent it perceptually"
Just to be clear that the statement applies to any releases meeting the A) criteria, not just 44.1 kHz @ 16-bit ones.
https://video.xiph.org/vid2.shtml
or on YT if you can't play it https://www.youtube.com/watch?v=cIQ9IXSUzuM
96kHz was created to better reproduce 20kHz high frequency, so the digital noise shaping filter does not need to be super sharp right at the Nyquist frequency.
Both were introduced for a sound technical reason. beyond that, most are marketing non-sense to cheat consumers.
[1]: https://www.cnn.com/2023/03/30/world/plants-make-sounds-scn
It’s like having gigabit internet to my house: I don’t actually need it, but when a website is slow, I know the problem isn’t in my internet connection.
I opened a support ticket but they never responded. After that it was difficult to take their lossless claims seriously when the labels were providing such garbage source material. Their whole value prop was totally hollowed out.
I don't know whether the labels still impose such horrible practices, but I largely gave up on streaming services after that experience and now focus on keeping good digital archives of my physical library.
I’ve played with the nice toys, and they are nice, but for 100x the price, they barely deliver 1.5x the experience.
https://www.carwow.co.uk/blog/carwow-quarter-mile-400-metre-...
https://en.wikipedia.org/wiki/List_of_N%C3%BCrburgring_Nords...
My favourite: "audiophile-grade" audio players which allocate a single continuous buffer of RAM into which they load/decode the whole .WAV/.FLAC file, because supposedly the CPU "jumping" between "fragmented audio" causes audible "jitter".
Of course, they don't know that what looks like continuous memory to user-code is probably discontinuous in kernel/physical RAM.
Didn't check in many years, I wonder if they created kernel level players to account for that, to have "true continuous memory"
Thanks for the laugh... this is absolutely bonkers. In case anyone is wondering, before sound hits our ears it has to go through a digital to analog conversion, which takes place on hardware independent of the CPU, operating with its own clock and buffers etc.
It gobbled like 90% of the CPU and I had to make sure I gave it a pretty large buffer so it didn't stutter when an app claimed CPU for more than a second, but it worked.
Also my headphones are extremely sensitive. I can touch the ring and sleeve of a jack with a finger, and touch a metal bed frame with a tip and I hear quiet clicks as I move the tip along the metal. Sometimes I do not even need to touch the jack with a finger. It doesn't work with small objects like a knife though.
And this is not the case of 'stronger than', it's 'strong enough to be caught up by anything resembling an antenna'.
You hear the interference because some analogue tract of your system:
Do transmit the radio waves everywhere
Does receive the inference from the all electric things around.
And now I'm skipping 'cuz I'm inna a bar and it's more interesting.
How this would occur without also producing grossly audible pitch distortion never seems to be discussed.
I use microphones that can 'hear' up to 100kHz (Sanken CUX100K) and for film sound design playing 192kHz audio at half and quarter speed the results are very significant, and reveal there IS 'content' above human hearing. Irrelevant for general listening but very important for sound design.
Nobody uses 32 bit float for recording (to do so is just to capture at least 10 bits of noise, most of that being brownian); its strictly a format for mixing and processing. You don't get any more resolution from 32 bit floating point than you do from 24 bit integer formats, but the result of "clipping" is less dramatic, hence the appeal of the format.
While there is some evidence that non-auditory human sensory perception may be sensitive to ultrasonic acoustic waves, it's pretty weak right now, and somewhat in the "woo" zone. It may turn out to be significant, or it may not. I wouldn't base an audio production workflow that requires 4x the cpu power and 4x the disk space on such tentative claims, but you're welcome to.
Yes they do, almost all high end field recorders used for film work are 32-bits now and have been for much of the last decade, often with some fancy preamp integration so that there is no expertise required for gain staging the recording. (I believe the implementations use a second matched 24bit ADC with 48 dB less gain in front of it).
The result obviously doesn't have a noise floor which is lower (as the noise of a room temperature _resistor_ gets in the way of that even at the 24-bit level) but they have more dynamic range so that your recording isn't ruined by hard clipping some unexpected loud sound.
It's a big improvement for practical usage, and also likely does improve SNR somewhat because you can run higher gains without as much fear that you'll ruin the recording. The reason it would pay off is that the SNR loss you get from splitting the signal is easily smaller than the SNR loss you would get from gain reduction to avoid clipping.
(maybe... capsule self noise is also limiting... at these levels, and usually people aren't using microphones designed for the lowest possible self noise unless they're doing something special)
There are ADCs that will provide 32 bits per sample but that's entirely different.
Current technology limits the bit depth to 18-22 bits and going beyond that you'd be very quickly recording brownian (atomic) noise anyway.
The point about 32 bit float is that it is a useful format for mixing, editing and general processing, so it is widely used in digital audio tools. But it is not a format that ADCs generate "natively" via their electronics - almost all of them are generate a 24 bit integer or fixed point value and then just supplying that as a 32 bit float value because the software asked for it (the software could have done it all by itself.
[EDITED: DAC->ADC since that is what I meant and what this is all about]
so maybe they do sample at 24 bit at a well chosen gain level and then convert to 32 bit float, with the max 24 bit value being above 1.0 float
or as GP said, use two separate ADCs at two different gains and combine their output
Of course it does! And that's what it does, of course. But that has absolutely nothing to do with the AD process itself, which is chip-limited to 24 bits and likely physics-limited to somewhat less than that.
You can't beat the physical limit of a DA circuit by doubling them up at different gains.
And .. you don't want to. Going beyond 22 bits gets you into brownian noise pretty quickly, which is completely pointless.
The best you can do (or could do) is get a very, very, very good DA that can really do 22 bits (likely not commercially available because of the expense), and then get the samples from it in whatever format works best for your purpose (24 bit integer, some fixed point value, or 32 bit floating point).
but what if you "allow" double that voltage and call it 2.0 float? a strong pressure into the microphone generates a stronger voltage
thermal noise limits you on the quiet signals, but not on the powerfull ones
so 22 bit for typical -1.0 -> 1.0 range and you can add a few more bits on top of that for stronger audio pressures (voltages) which you would traditionally clip
This is a marketing page after all, explaining point-wise samples, reconstruction filters and "staircases" is way beyond scope.
Same as in images, pixels are not "little squares"
You have some low noise amplifier. There is a signal. You split it. The result on each side has >=1 bit worse noise floor, probably somewhat worse as we're not using superconductors :P-- as you expect: there is no free lunch.
Now: take one copy and attenuate it 48dB, further degrading its noise floor. Sample both. The attenuated copy is mostly useless, except when the input goes high enough that it would have hard clipped the other ADC.
So the tradeoff is that you lose a small amount of noise floor constantly-- out at the 20th bit, that you probably didn't care about (microphone self-noise is limiting you out there anyways at normal volume levels), in exchange for never clipping.
To turn this into a better ADC generally, you'd need the splitting stage to not hurt the noise floor, but it does.
The reason it's not the same as just lowering the gain so that you won't ever clip is that to get the same dynamic range you'd have to lower it by 48dB and now your ADC doesn't achieve its potential for typical signals. You could lower the gain by 3dB (or whatever the splitting cost you) and get the same results for the low gain signal and a little more headroom, but you would not get the massive headroom increase of this approach.
For this to work one must also have amplifiers with much wider dynamic range and SNR than ADCs, but we do.
The natural output for this approach is a float-- the most natural would be a weird float where instead of an exponent one bit tells you which ADC is in use and represents a factor of 256 or whatever, but in practice these recorders just output 32-bit floats. I haven't looked but I wouldn't be surprised if there were only two exponent values ever used in their output.
So, basically, no better than the best AD converters we already have?
My understanding of the fundamental limit to AD performance is that the brownian noise level is around the 22nd bit level. So even if you come up with techniques to successfully measure down to that level, you're basically picking up .. inevitable, irremovable, irrelevant noise.
Possibly there are gains to be made by not worrying about the noise floor and caring more about the lack of clipping, but I'm not seeing people screaming about that. The "noise" seems to be "N bits of dynamic range", not "slightly less dynamic range but it will never clip!"
A common experience for someone doing field recording of performers (my experience is music) is you twiddle your setup to get the gains reasonably high to get good SNR even for quiet parts. ... and then you record the actual performance, and you find that the tuba player really got into it for the real performance and the new peaks are 10dB over where they were in the practice. And now your recording is screwed up with a bunch of hard clipping you have to deal with. So then experience tells you in the future to take whatever you thought was safe and lower gains another dozen db.
The multi-ranged recorders eliminate that problem and the result is that you don't need to use precautionary gains, and you get a better SNR in your recordings. You probably don't need to adjust gains at all: The gain can be whatever makes the self-noise of the microphone dominate the SNR of the process, ... which would be too high for the loudest samples, but the clipping handling deals with that.
The samples that need to use the extended range have worse SNR (and probably poor linearity due to mismatches between the converters), but human hearing is much less critical to noise with loud signals anyways.
That's what could be done if ADCs were perfectly linear and noise free and limited only by their bit-width. Sadly, they are not. The non-linearity one can in theory measure and correct for, but the noise can be corrected for only by oversampling. And then you might as well use a single ADC of lesser bit width and higher sampling rate.
If, say, two 24b ADC (20b noise free, non-linearity 2LSB) with one receiving the input signal with an approximate 10bit higher gain (+60dB) and one would combine their outputs with that 10b shift (and ignoring the input of the low gain path, if the signal falls below a given threshold to reduce the noise contribution of that ADC and the input of the high gain path if the signal exceeds another threshold in order to avoid clipping), then one could construct a 32b float.
This doesn't improve resolution (which arguably would be pointless) or linearity (not all that critical in audio methinks) but dynamic range, which I can see some appeal of (in extreme recording situations, say you'd want to record the breathing of a shooter followed by the gun shot -- there remains the challenge of finding a microphone capable of a 120dB range, but perhaps one could use two different ones ...).
No one is arguing that there are practical audio microphones + ADCs that produce accurate, undistorted 32-bit float output across the full representable range. But they don’t need to! For professional use, the ability to produce perceptually accurate output, with inaudible noise, across a very wide dynamic range, is extremely useful. Think of it as fancy, real-time AGC. It does not need to be perfect. If you can record a loud transient without substantial distortion, and also record sounds with 2^16-fold lower amplitude (~96dB lower) while still remaining well above the noise floor immediately after the transient is gone, this ability is useful. Plenty of real-world noises are well above 120dB, and plenty of human-audible sounds are below 20dB. You can’t play back the recording, at least not without making parts inaudible or injuring your audience, but you can edit it. And a setup like this lets you do it with one microphone and no fiddling with gains in advance.
> Nobody uses 32 bit float for recording (to do so is just to capture at least 10 bits of noise, most of that being brownian);
This is not true and not true for a good and important reason! One which has no bearing on the kind of DACs that exist.
Modern field recorders allow gains set a 'reasonable' level that maximizes SNR for recordings but still won't clip when there are much louder peaks. Not so dissimilar to how a 6-digit multimeter can achieve its advertised performance both on a 0-5v range and a 0-300v range but cannot give more than 6 digits at the higher range.
Obviously, everyone and their mother uses 32 bit float as an internal sample format because of its fitness for purpose (except the folks who think they need 64 or 80 bit floating point, of course). But they are not using "32 bit floating point samples" - the samples come from an (at best) 18-22 bit integer conversion.
Due to their high cost such ADCs have no longer been used in audio for many decades. They may still be encountered in some expensive measurement instruments that need high resolutions at significantly higher sampling frequencies than needed for audio.
All audio ADCs have a very low resolution per sample, e.g. 4 bits or even lower, but they sample at a very high frequency, of many MHz. Then the bit stream is digitally processed to generate whatever format is desired for output, at a lower sampling frequency and a higher resolution, e.g. 24 bits @ 192 kHz.
There is a difference between the actual resolution at the output and the effective resolution, which is limited by noise, e.g. the 24 bit samples may have an effective resolution of 20 bits or 21 bits or 23 bits, etc., i.e. they contain noise with an amplitude corresponding to those effective resolutions.
The digital algorithm that converts the low resolution input samples (e.g. 4 bits @ 5 MHz) inside the ADC can easily be modified to generate a different numeric output format, e.g. FP32.
Neither FP32 nor 24-bit is the native format of the A/D conversion. If the ADC outputs FP32, that is even more convenient for further audio processing. Obviously, the quality of the ADC is independent of whether it outputs FP32, and the FP32 samples will have a different effective resolution on each ADC, which seldom would be as high as 24 bits, due to the noise.
> There are ADCs that will provide 32 bits per sample but that's entirely different.
Now that requires elaboration.
There is e.g. AD's LTC2500 (https://www.analog.com/en/products/ltc2500-32.html). Not meant for audio (too slow at 32b) and not noise free, but it's a bona-fide 32b ADC.
Now there might be no ADC which provides 32b wide noise-free samples at sample rates needed for audio and given the absurdly low level of a LSB signal that might be as infeasible as it would be pointless, but that's a bit of a different statement.
Surely you understand a recording made at 48kHz has a max freq response of 24kHz and played at half speed that max freq is 12kHz and at quarter speed only 6kHz. You can very clearly hear the filter cut off due to Nyquist. Record at 192kHz with mics capable of 100kHz capture and when played at quarter speed, the sound is full spectrum because there is no truncated frequency response. And when I load a 192kHz recording to izotope RX I can literallu see the harmonics going up to 96kHz. (not with every sound of course)
I repeat, i am not talking about 'normal' listening. I am talking about an industruy you have no knowledge or lived experience with, so spare me the incorrect claims about what can & cant be heard.
I'm the original/lead developer of Ardour, a cross-platform DAW, and have been working with digital audio for more than 25 years.
There are no 32 bit ADCs - your SD MixPre's are giving you (at best) 22 bits packaged as a 32 bit float value. The preamps make absolutely zero difference to the AD conversion (though they might sound real nice).
> Surely you understand a recording made at 48kHz has a max freq response of 24kHz and played at half speed that max freq is 12kHz
This is a very naive version of what "played at half speed" might actually mean. If properly and correctly resampled, this is not true.
> And when I load a 192kHz recording to izotope RX I can literallu see the harmonics going up to 96kHz
Well, I'd certainly hope so! But the question is: what are the energy levels associated with the partials above Nyquist? If you recorded at 384kHz with sensitive enough equipment, you'd see partials above 96kHz - but at extremely low energies because ... well, that's just how physics works.
[EDITED to remove AD/DA confusion]
The half speed you call naive is again just showing your ignorance. Sound editors have been using this technique since the days of recording on a Nagra at 15ips and literally replaying at 7.5ips half speed, and at 3.75ips for quarter speed. There is nothing naive about it, it is a very well know technique. To be able to achieve the same result digitally with full spectrum has impacted every feature film you have experienced in recent years. Again I speak from decades of lived experience.
My use of DAC was a thinko, I've edited at least post to correct it since in the current context we're always talking about ADC. Apologies for that.
You are arguing about techniques you have no experience with.
I am extremely aware that as a data format in DAWs and other recorders, 32 bit floating point is completely common.
For the same reason, video processing is preferably done on FP16 samples of the color components even if both the input ADCs and the output video signal may use only 10-bit or 12-bit per sample, at most.
Moreover, most high-resolution audio ADCs do not really sample the input audio at a 24-bit resolution, but they use only a sigma-delta method where the actual samples have only a few bits, possibly only even 1 bit.
Then DSP techniques are used to convert the audio stream with a high sampling frequency and a low resolution per sample into an audio stream with a low sampling frequency and a high resolution per sample, which is the external output of the ADC.
If you had access to the raw audio bit stream as actually captured by the ADC, you could modify the decimation algorithm to really output FP32 samples, though no existent ADC could actually have a so high dynamic range (except if the output bandwidth would be reduced a lot, to filter the input noise).
Or, in some cases, FP64.
It kind of changed me a bit when I ran through 20 lossless tracks I had re-encoded to various mp3 bitrates and realized that even on a fancy system, it can be really hard if not impossible to discern even moderate lossy from lossless.
If you are an audiophile geek, really think about if you want to try this, the reality check might crack your foundations.
We store files in the highest quality because it gives us the option to encode the music without audible loss of quality.
that's a bitrate of 1GB per 9-12h, and for cases when it's too much I just have cached music on my device (I'm lucky to have mostly empty storage on my 256gb phone)
In a nutshell: nullc, rahimnathwani, zamadatix and vor_ know their shit, and geraldmcboing and PaulDavis are technicially correct but talking past each other. speak_on and TheOtherHobbes are confidently wrong.
And also: 44.1 kHz captures the entire human audible spectrum with room to spare, and 16-bit already goes beyond anything useful for listening. The higher resolution / sample size format is useful for production or archival purposes only.
The two main reasons why you hear a difference between the two formats: (1) it's likely a different master, (2) tiny gain differences in the signal (salesmen use this trick, but it's also easy to do it by mistake).
Most modern digital synths have already caught onto this and run internally at much higher sampling rates even if their output gets downsampled, but sometimes you run across a vintage plugin that runs at the host audio rate and working in a higher sampling rate is audible.
Oversampling gives you headroom for aliases for the rest of the synth that is more vulnerable to it.
* Some people are still making this mistake, despite information on the (many) ways to do it the right way being widely and freely available!
There's some ways to do band-limited distortion but...they aren't nearly as widespread, easy, or universal as band-limited oscillators.
Ring modulation is funny though because you'd ideally want the sidebands to modulate down by default rather than filter them out, that's why you're using it.
So if your synthesizers do not use proper band-limited oscillators then 192KHz is _FAR_ too slow. You'd want to be running at hundreds of KHz, perhaps a few MHz.
In reality synth software that doesn't sound like crap uses band limited oscillators and should work okay at 48KHz too. That said, even if the oscillators are band limited it may be the case the varrious modulations aren't band limited properly, as getting those wrong won't sound instantly wrong (in particular because you have to modulate to make it wrong, and the underlying change of the modulation may make it harder to tell its wrong).
Though also in those cases if you're not counting on every step being properly band limited then 192KHz may be an improvement but you're still probably getting some meaningful aliasing. I think given how fast computers have become relative to digital audio there is probably a good case to just make any "modular synth" run at 32-bit 480KHz or even 4.8MHz through every stage that could process the audio.
Maybe 192KHz really is enough to suppress the aliasing artifacts but I think to be convinced of that I'd want to see a system that supported both and validate that the difference between a downsampled 48KHz output from the two modes was below -90dB or something.
Or otherwise you can just declare that the aliasing is part of the sound and then there are no right choices... 24khz sampling, 48k, 192k ... who cares, use what you like best. :)
1. It should run at FP64 if you want to preserve filter resonances, etc.
2. At 10x/100x fixed-rate oversampling, even a modern "fast" CPU will have very few cycles per (higher-rate) sample to run the DSP for 1 "module" of the software modular. Forget about interconnected modules, multiple tracks, or polyphony. For this kind of "analog"-style processing, it's better to run adaptive-rate algorithms (think SPICE) instead of wasting compute on unnecessary extra audio samples.
For adaptive rate I think the issue there is you have a hard-realtime constraint for this usage (even if you wouldn't mind rendering offline, you kinda have to hear it realtime to tweak it-- after all you might tweak it in a way that brings out an artifact you like and then be disappointed by the render). Also in the case of a whole modular system having all sorts of different parts needing to be part of the adaptation loop seems pretty hard to me.
My thinking was just in general that 192k is really not enough to prevent aliasy algorithms from messing up. If you are alias safe you can probably run at 48k and be fine. If you're not, you really want to go much higher.
These simulations are single core to avoid core-to-core latency. Number of cores isn't relevant unless you want to run independent voices/channels and sum them at the end.
So you start with a very optimistic ~90 GFLOPs of 64-bit FMA on Zen 4. Unfortunately, not all operations are clean multiply-adds. Realistically, you'll need trigonometric functions and LUTs, which are quite slower. Btw, the tradeoff between when to compute vs LUT is very fragile and can change due to a ton of factors (notably integrator algorithm).
Then the data you are operating on won't fit cleanly in AVX-512 registers, requiring spills to L1 cache. Ok, still fast on a modern core.
Of course, the peak theoretical number assumes clean vectorization with double-pumped AVX-512... which also won't happen in practice. Classical DSP will fare better (https://www.youtube.com/watch?v=Ssq0a-YdamM) but SPICE integrators are inherently branchy and divergent. Especially for adaptive integrators, you'll waste a lot of operations trying to "lock in" at the exact time point where the waveform turns a corner. Apple Silicon is better at this messy, branchy code.
So yeah, it's possible-but-hard to hit hard-realtime under these conditions.
192 for mixing and mastering can be useful especially if you're doing a lot of effects, especially anything that pitch shifts. But I've seen low quality phone-microphone recordings make it to the master; if you capture lightning in a bottle, it hardly matters what the settings were, what the microphone was, or anything else.
Some previous discussions:
2023 https://news.ycombinator.com/item?id=34698427
2022 https://news.ycombinator.com/item?id=30138561
2019 https://news.ycombinator.com/item?id=19318898
2017 https://news.ycombinator.com/item?id=15127633
2015 https://news.ycombinator.com/item?id=10520639
And its all good! It's perfectly fine to say "I prefer the sound when the whole mix (or just that guitar) ends up being subject to interesting and possibly harmonically relevant distortion at low levels".
Just don't say "The version with the distortion is more accurate than the one without", because that's a lie.
(2014) https://news.ycombinator.com/item?id=8689231 424 comments
(2015) https://news.ycombinator.com/item?id=10520639 228 comments
(2017) https://news.ycombinator.com/item?id=15127633 428 comments
(2019) https://news.ycombinator.com/item?id=19318898 314 comments
However, the article claims that the final distribution doesn’t need to have a bit depth of more than 16. That does not match my experience. I can tell the difference between my renders that are 16 bit vs 24 bit. I cannot tell the difference between 44.1 kHz and higher sample rates, and that’s consistent with the math (Nyquist-Shannon), but bit depth is a different matter. Would be fun to participate in a double-blind test that includes my own tracks and others.
I'm writing in jest, but long time ago, -hp- used actively cooled FETs (not a very popular approach today as that caused problems with condensation and we have better FETs now).
established using double blind testing, I assume?
You can the focus on other things.
Example: I Bought the best skis possible. Now I know I need to just focus on my skills and not blame the equipment.
The problem is the people spreading myths and disinformation out of ignorance or to promote their enterprise.
The weak links are producers/mastering-engineers, speakers/headphones and the room when using speakers.
As for how this relates to audio compression, in particular in the context of 2012. you are making a tradeoff of storage size and decompression cost. Maybe that doesn't matter to you, but maybe it either did in 2012 or still does.
And none of them are broken after 20 days unless it's low tide or I fuck up on a cliff band. I ski a minimum of 100 days a year, and the only thing you can notice after 60 days is some slight decambering on softer skis.
I will say that boots tend to get soft around 100 days. But usually dealing with that's also a skill issue; get good balance, and you don't have to have the boot hold your sorry form in place. It's how people basically skied for a thousand years with leather boots: They were good.
OTOH, we know nothing of your audio equipment nor how its setup.
EDIT: Did some more ABX testing with a CD-quality track that I'm much more familiar with ('Introduction' from the Mirrors Edge soundtrack, which has been my go-to for comparing audio gear for the last decade). I could sometimes distinguish 128k mp3 this time, though interestingly, I got it consistently wrong rather than right. For some reason the compressed version seems to be my preference. Dropping to 96k mp3, I got it right 100% of the time - though only because there was a very noticeable difference in the stereo positioning of the first sound, rather than a difference in the quality of the sound. I think if it were mono I would still be unable to tell.
At 48khz, with a tone generator (not rando Youtube videos or anything), you should be able to clearly hear up to around 16khz (as in, can tell pitch of tone), and be able to tell a tone is being played at all up through 18khz, and hopefully up to 20khz.
It’s like photographers who are confused about the difference between raw and bitmap (jpeg), videographers confused about the difference between linear raw vs log vs gamma encoded, etc.
Just because a data format with higher bit depth/sampling frequency/whatever exists for editing purposes, doesn’t mean it’s “better” or makes sense as a consumption format for a finished work.
Forms of manipulation bring inaudible content into the audible range.
Of course that doesn't mean audiophiles aren't being audiofooled by it, but there is legitimate usage.
I like your point about editing, because trying to mix too much tracks in 44100 or 48000 makes spontaneus click sounds which are not supposed to be here.
To try to imagine something similar: the human eye is unable to see UV light, yet fluorescent paint has a visible quality of its own compared to "normal" pigments.
this has practical applications
I use a DAC by focusrite which can do 24-bit, and if I want to listen to higher fidelity audio on my planer headphones then I should be able to. Why should I limit myself to 16-bit
If I like an artist that I find on streaming, I buy an LP and get a lossless download for free. I still have a music library and I will never rent my favorite music.
Artists prefer to connect directly with their fans and BC is probably the best platform for people who care to pay and support acts directly. They have high res downloads and I import them.
Also the playback rate and the file rate are different topics. The former can get into scenarios more like the audio processing section of the article e.g. I had this one shitty headset for work which required me to set the volume to 1-2 (out of 100) on the computer and I could actually blind test tell when it was in 16 bit or 24 bit mode because it was cutting and boosting it so much it effectively lost precision in 16 bit mode.
I can always tell if my 44.1 songs are being resampled to 48 because they're being run through the OS mixer
But a quality audio player should account for this and do it's own.
It is an incredible resource to see the quality of the resampling algorithms used by the actual production software likely used in any digital audio workflow.
You will see that while the best are indeed almost 100% transparent, many are not.
your software is among the best, but not pitch black best :)
There is also https://src.hydrogenaudio.org/ (with no IP based restrictions, AFAIK).
There's multiple YouTube channels that I listen to as podcasts, that are professionally created and the creators presume that exported audio works like studio audio, so what you end up with is really quiet audio that can't be turned up without pre-processing.
If we distributed audio the same way we work with it in a studio, we could forgo a lot of problems.
Also, the human ear does have enough dynamic range to make 24 bits worthwhile, though that much dynamic range is rarely used in recordings, and that high of a bit depth provides no benefits within a small dynamic range. A 192 kHz sample rate, on the other hand, is always useless.
https://www.routledge.com/Sound-Reproduction-The-Acoustics-a...
And a talk by the author covering some of the material for those who prefer that: https://www.youtube.com/watch?v=zrpUDuUtxPM
It would have cost the same for the entire stack to be 16bit/44.1kHz at every step, but with excessive resolution I can control the volume anywhere. The bits right before the analog conversion at the end are essentially the same whether I turn down the volume in the software player, the operating system, or the DAC/amplifier.
When I play from the computer, I'm not sure whether it is using the clock on my Mac, the clock on the optical interface, or the WiiM's clock. However, I do not notice any difference in fidelity when I use the Qobuz software player on my Mac or use Qobuz Connect to allow the player to directly stream from the source, so either it isn't a difference that I can hear, or the WiiM's internal clock is used for both sources.
https://ardour.org/ is my website.
Firstly, it's an amazing experience to randomly interact with people like you - I love and use your software. Hats off and thanks for what you offered to the industry!
But secondly, your statement makes even less sense to me: obviously artifacts do add up. Yes, not linearly, like any complex audio in general. But the more tracks with artifacts I have, the more artifacts I have overall. It's not like they cancel each other (outside of normal frequency cancellation).
The human threshold-of-hearing curve intersects the threshold-of-pain curve at about 20 kHz.
Above that frequency (or thereabouts) the sound has to be so loud that it will literally instantly damage your hearing before you can hear it.
This has been replicated across many studies for more than 100 years.
Flicker threshold is completely different. You can’t damage your vision by increasing the FPS, and it has always been commercially desirable to use a lower frequency because that is cheaper.
In addition, nobody cares about "measurable" artifacts (or rather, they should not). What matters are "audible" artifacts. We have measuring equipment that is vastly more sensitive than human ears (e.g. your recording equipment that can pick up signals far above 22kHz). What's measurable is not particularly interesting - what's audible is.
Artifacts do not sum linearly, because they do not originate from correlated sources (unless you're doing something rather unusual).
Glad you can hear the difference between two converters, but I trust you've tested it in a double blind setting?
And absolutely - I blind tested coverters extensively. Mbox2, Black Lion Audio upgraded converters, UA, Prism.
Yes, the discussion was "never about analog vs AD". But my point is that I see little point wasting time on one set of artifacts (in the digital realm) that are tiny compared to those introduced in the analog realm. If there's a mouse and an elephant about to enter your home, you focus on the elephant, no?
The big difference, of course, is that "everyone" has convinced themselves that most/all of the analog artifacts, as big as they are, are somehow "tasteful" or "artistic", whereas the digital ones are just "math errors". I don't think is too helpful.
And look, if lots of people could get through double blind tests and still show they can hear aliasing or whatever the digital artifact du jour is, then I'd say "yes, absolutely, we need to be very aware of this and do everything we can to reduce or eliminate it". But as far as I can tell, this just isn't the case.
To your main point: yes, all artifacts are just our learned, cultural, developed preferences. In the exact same way major/minor thirds were considered dissonant just a few hundred years ago - it's all a learned perception, not an absolute judgment.
I would go even further, doesn't matter whether people perceive aliasing as a major issue, it's no different from the U47 "warmth". You can't afford this, probably, as a software developer in a way, but at the most fundamental level any sound's - or artifact's - judgment is based on our our current diagram of "sounds nice" vs "sounds bad".
Who has the best ears? What can they detect?
I know from my 20-ish year mixing experience that I can hear the difference when mixing. Is it good evidence? No. So we can agree to disagree then.
I'm not disagreeing with you. I'm really curious about the limits of what people can hear, what can be taught and what is rare.
Now in terms of realistic audio encoding, 16 bit at 44.1 kHz is designed to be a faithful representation as far as human hearing is concerned. Can someone with a trained ear potentially tell the difference between that and 24 bit at 192 kHz? In a studio environment it's possible. Most audiophile claims are dubious and a blind A/B test catches them out on most of it but the Nyquist-Shannon sampling theorem does not directly apply to quantized samples, it's about exact samples and with quantization, sampling rate is intertwined somewhat with the quantization depth.
A quick search returned this PDF with a nice diagram of what aliasing looks like: https://download.tek.com/document/76W_30631_0_HR_Letter.pdf
To draw a design parallel: pixel-perfect design isn't something we are born with, noticing tiny details is a developed skill.
And yes, you are on point: oversampling is used extensively, but this just points at the exact issue: Nyquist theorem gave us a math algorithm, we still need to account for the electronic component imperfections. And then we are entering a different space of quality/precision/psychoacoustics/perception/etc. Meaning, not all converters, not all pre-amps, not all mics "sound" the same, even when they use same types of components on paper.
Do you have more convincing sources?
Would be happy to see an actual, real study to prove that humans can notice, but to my knowledge none exist that confirm they can. Not even any on teenagers or younger (the only group that can even hear close up 20khz).
The energy of the signal components above the Nyquist is generally very low, and very few double blind tests have given any indication that humans can detect the resulting aliasing (even though many people claim to be able to do, almost always in non-double-blind environments).
Badly written digital synthesis can generate high energy signal components above 22kHz, but that's because they're badly written, not because the theory is wrong.
This space is not driven by a single precise formula. 48/96 kHz helps some engineers to produce better sounding mixes. Can everyone hear the extended range of Adam tweeters? Probably not. But some can, and they benefit from that. Even if there is no double-blind study to prove this in absolute terms.
But very little music is like that, and the energy profile above Nyquist will differ dramatically. Consequently, you're not summing a set of identical aliasing results, and in general, the results will still be undetectable to almost everyone.
Jacob Collier routinely works with 300+ tracks in Logic. He doesn't worry about this sort of thing, and neither do the Grammy voters who love what he does.
It is always amazing how much that is claimed about what people can hear fails to show up when tested in this, the only acceptable scientific way.
Perhaps Maserati has done this, and could still tell the difference. In which case, he should carry on! But he should carry on anyway! People should do what brings them joy, and if he likes working at 44.1kHz or whatever, he should absolutely do that.
What people should not do is lecture about stuff that isn't true and/or isn't demonstrable in proper test settings, and most (not all, but most) of the SR stuff fits into one or other or both of those categories.
And since there are no double-blind studies supporting this tech, both using and adding any of these features to your software would only be propagating this scam further? I.e. far more than just lecturing "about stuff that isn't true" this actually physically implements features that are not true?.. OK, at least you are staying consistent.
(Also, looking forward to you discovering that there are not many double-blind studies supporting the "delusion" that the effect of a low pass filter is real and not just something we can measure but can't hear.)
There are places where (a) double precision (or better) floating point math benefits DSP, but that's nowhere in Ardour (and likely, if one is clear about the definition, in any other DAW either) - certain types of plugins can benefit from this, but they should not impose that cost on the rest of the processing infrastructure (i.e. they should convert internally and then back again, rather than require that the whole host uses 64 or 80 bit floating point.
There is a reasonably good argument for 96kHz because of the possible characteristics of the brickwall filter and its impact on aliasing. This argument gets a lot weaker for even high SRs.
However, both of these are largely theoretical in the sense that very few people, if any, can reliably hear the results in actual produced music or soundtracks. So while it might be a good idea to use higher SRs, whether that actually results in something that can truly be said to "sound better" is much more questionable.
At the highest level, I think that all of this is pretty irrelevant. Most of the music that people consider "great" was recorded with equipment far below the quality levels achievable with mid-priced pro-sumer gear today (mics perhaps being the sole exception). Great music/great sound is, I think, appreciated largely independently of its audio "fidelity". There is a huge difference between an amazingly well-recorded ensemble and a poorly recorded one, but for most people, if the well-recorded ensemble is playing stuff they don't like and the poorly recorded one is playing some stuff they love, all the "fidelity" in the world won't improve the former and the lack of it won't stop them loving the latter.
In addition, with the rise of deeply impressive sample libraries and better and better synthesizers running as plugins, the AD/DA elements of music during "recording"/composition are becoming less and less important for more and more music. That amazing patch someone uses in Onmisphere doesn't get better or worse by using a higher SR, and that incredible string library (e.g. New Albion) is what it is regardless of what your converters can do.
So, in short, I would say: do what you want, but if you're going to try to justify it with science, make sure the science is right and if you're going to try to justify it with "sound", then be aware that for most people the differences won't matter (or even exist at all).
Also, sadly consumers are getting used to low quality audio nowadays - they often listen to lossly compressed audio on social media (sometimes decompressed and re-compressed several times) which is then re-compressed to send to bluetooth headphones, or played back on an awful smartphone speakers. Streaming services also use compressed audio.
So I guess the programmer equivalent is distributing .pdb's (or, symbols)
I'm not interested in finetuning everything in my life for efficiency.
in DSP it matters a lot, so if mixing digitally or producting remixes etc. its useful to have more and larger samples to work with
It's a bit redundant for a skilled technician, they're already used to setting the gain staging, inbound compression, and feathering the mics to avoid this in 24-bit, but if you're handing a boom mic to a novice and have a scene where e.g. someone's whispering and another person's screaming, it can be nice to not have to worry about it.
Don't forget to buy the new low oxygen platinum plated HDMI cables for the better experience!
/s