https://ardour.org/ is my website.
Firstly, it's an amazing experience to randomly interact with people like you - I love and use your software. Hats off and thanks for what you offered to the industry!
But secondly, your statement makes even less sense to me: obviously artifacts do add up. Yes, not linearly, like any complex audio in general. But the more tracks with artifacts I have, the more artifacts I have overall. It's not like they cancel each other (outside of normal frequency cancellation).
The human threshold-of-hearing curve intersects the threshold-of-pain curve at about 20 kHz.
Above that frequency (or thereabouts) the sound has to be so loud that it will literally instantly damage your hearing before you can hear it.
This has been replicated across many studies for more than 100 years.
Flicker threshold is completely different. You can’t damage your vision by increasing the FPS, and it has always been commercially desirable to use a lower frequency because that is cheaper.
In addition, nobody cares about "measurable" artifacts (or rather, they should not). What matters are "audible" artifacts. We have measuring equipment that is vastly more sensitive than human ears (e.g. your recording equipment that can pick up signals far above 22kHz). What's measurable is not particularly interesting - what's audible is.
Artifacts do not sum linearly, because they do not originate from correlated sources (unless you're doing something rather unusual).
Glad you can hear the difference between two converters, but I trust you've tested it in a double blind setting?
And absolutely - I blind tested coverters extensively. Mbox2, Black Lion Audio upgraded converters, UA, Prism.
Yes, the discussion was "never about analog vs AD". But my point is that I see little point wasting time on one set of artifacts (in the digital realm) that are tiny compared to those introduced in the analog realm. If there's a mouse and an elephant about to enter your home, you focus on the elephant, no?
The big difference, of course, is that "everyone" has convinced themselves that most/all of the analog artifacts, as big as they are, are somehow "tasteful" or "artistic", whereas the digital ones are just "math errors". I don't think is too helpful.
And look, if lots of people could get through double blind tests and still show they can hear aliasing or whatever the digital artifact du jour is, then I'd say "yes, absolutely, we need to be very aware of this and do everything we can to reduce or eliminate it". But as far as I can tell, this just isn't the case.
To your main point: yes, all artifacts are just our learned, cultural, developed preferences. In the exact same way major/minor thirds were considered dissonant just a few hundred years ago - it's all a learned perception, not an absolute judgment.
I would go even further, doesn't matter whether people perceive aliasing as a major issue, it's no different from the U47 "warmth". You can't afford this, probably, as a software developer in a way, but at the most fundamental level any sound's - or artifact's - judgment is based on our our current diagram of "sounds nice" vs "sounds bad".
Who has the best ears? What can they detect?
I know from my 20-ish year mixing experience that I can hear the difference when mixing. Is it good evidence? No. So we can agree to disagree then.
I'm not disagreeing with you. I'm really curious about the limits of what people can hear, what can be taught and what is rare.
Now in terms of realistic audio encoding, 16 bit at 44.1 kHz is designed to be a faithful representation as far as human hearing is concerned. Can someone with a trained ear potentially tell the difference between that and 24 bit at 192 kHz? In a studio environment it's possible. Most audiophile claims are dubious and a blind A/B test catches them out on most of it but the Nyquist-Shannon sampling theorem does not directly apply to quantized samples, it's about exact samples and with quantization, sampling rate is intertwined somewhat with the quantization depth.
A quick search returned this PDF with a nice diagram of what aliasing looks like: https://download.tek.com/document/76W_30631_0_HR_Letter.pdf
To draw a design parallel: pixel-perfect design isn't something we are born with, noticing tiny details is a developed skill.
And yes, you are on point: oversampling is used extensively, but this just points at the exact issue: Nyquist theorem gave us a math algorithm, we still need to account for the electronic component imperfections. And then we are entering a different space of quality/precision/psychoacoustics/perception/etc. Meaning, not all converters, not all pre-amps, not all mics "sound" the same, even when they use same types of components on paper.
Do you have more convincing sources?
Would be happy to see an actual, real study to prove that humans can notice, but to my knowledge none exist that confirm they can. Not even any on teenagers or younger (the only group that can even hear close up 20khz).
The energy of the signal components above the Nyquist is generally very low, and very few double blind tests have given any indication that humans can detect the resulting aliasing (even though many people claim to be able to do, almost always in non-double-blind environments).
Badly written digital synthesis can generate high energy signal components above 22kHz, but that's because they're badly written, not because the theory is wrong.
This space is not driven by a single precise formula. 48/96 kHz helps some engineers to produce better sounding mixes. Can everyone hear the extended range of Adam tweeters? Probably not. But some can, and they benefit from that. Even if there is no double-blind study to prove this in absolute terms.
But very little music is like that, and the energy profile above Nyquist will differ dramatically. Consequently, you're not summing a set of identical aliasing results, and in general, the results will still be undetectable to almost everyone.
Jacob Collier routinely works with 300+ tracks in Logic. He doesn't worry about this sort of thing, and neither do the Grammy voters who love what he does.
It is always amazing how much that is claimed about what people can hear fails to show up when tested in this, the only acceptable scientific way.
Perhaps Maserati has done this, and could still tell the difference. In which case, he should carry on! But he should carry on anyway! People should do what brings them joy, and if he likes working at 44.1kHz or whatever, he should absolutely do that.
What people should not do is lecture about stuff that isn't true and/or isn't demonstrable in proper test settings, and most (not all, but most) of the SR stuff fits into one or other or both of those categories.
And since there are no double-blind studies supporting this tech, both using and adding any of these features to your software would only be propagating this scam further? I.e. far more than just lecturing "about stuff that isn't true" this actually physically implements features that are not true?.. OK, at least you are staying consistent.
(Also, looking forward to you discovering that there are not many double-blind studies supporting the "delusion" that the effect of a low pass filter is real and not just something we can measure but can't hear.)
There are places where (a) double precision (or better) floating point math benefits DSP, but that's nowhere in Ardour (and likely, if one is clear about the definition, in any other DAW either) - certain types of plugins can benefit from this, but they should not impose that cost on the rest of the processing infrastructure (i.e. they should convert internally and then back again, rather than require that the whole host uses 64 or 80 bit floating point.
There is a reasonably good argument for 96kHz because of the possible characteristics of the brickwall filter and its impact on aliasing. This argument gets a lot weaker for even high SRs.
However, both of these are largely theoretical in the sense that very few people, if any, can reliably hear the results in actual produced music or soundtracks. So while it might be a good idea to use higher SRs, whether that actually results in something that can truly be said to "sound better" is much more questionable.
At the highest level, I think that all of this is pretty irrelevant. Most of the music that people consider "great" was recorded with equipment far below the quality levels achievable with mid-priced pro-sumer gear today (mics perhaps being the sole exception). Great music/great sound is, I think, appreciated largely independently of its audio "fidelity". There is a huge difference between an amazingly well-recorded ensemble and a poorly recorded one, but for most people, if the well-recorded ensemble is playing stuff they don't like and the poorly recorded one is playing some stuff they love, all the "fidelity" in the world won't improve the former and the lack of it won't stop them loving the latter.
In addition, with the rise of deeply impressive sample libraries and better and better synthesizers running as plugins, the AD/DA elements of music during "recording"/composition are becoming less and less important for more and more music. That amazing patch someone uses in Onmisphere doesn't get better or worse by using a higher SR, and that incredible string library (e.g. New Albion) is what it is regardless of what your converters can do.
So, in short, I would say: do what you want, but if you're going to try to justify it with science, make sure the science is right and if you're going to try to justify it with "sound", then be aware that for most people the differences won't matter (or even exist at all).