24-bit samples is ridiculous overkill. That's a huge dynamic range that's completely unnecessary.
At 192KHz you'd be able to capture 96KHz signals, far, FAR outside the range of human hearing. Human hearing peaks at 22KHz so you only need a sample rate of 44KHz to capture the total range of human hearing.
For human voice you don't really need better than 16-bit samples at 12KHz or so. That's for great quality voice.
The only reason audio mastering is done at huge sample sizes and sampling frequencies is to prevent aliasing during mixing and to preserve higher frequency harmonics. There's absolutely no need for such rates delivering to human beings.
Also higher fidelity audio sampling is available for phone calls. The issue is more political than technical. Cellular carriers don't like to negotiate higher quality calls between one another so inter-carrier calls tend to fall back on the lowest common denominator AMR-NB codec. Intra-carrier calls don't even reliably pick AMR-WB let alone EVS available with VoLTE.
humans can’t hear above 20khz. adult humans can’t hear above 16khz or so, we lose the top end before age 20. this means that the standard 48khz sampling rate covers the entire human hearing range and then some (0-24khz). any sampling rate over 48khz for sound intended for human hearing is a total waste.
Also, you might possibly be sensitive to resampling artifacts if your output device runs at 44.1kHz and your file is 48kHz or vice versa.
Audio testing is hard, and testing on yourself is tricky... But if you have a sample that you're convinced sounds better at high rates than lower rates, I would urge you to put it through a tool to resample it down to lower rates and see if/when you can tell the difference. If the rate isn't an even multiple, it's worth using a tool that can dither; dithered resampling artifacts are less abrasive than undithered... I had some voice recordings to play over the phone, and everything needed to be 8kHz u-law; the 48kHz original recordings sounded better than 44.1kHz original recordings because one is even multiple and the other isn't, but either way, the waveforms looked worse than it sounded.
"Headroom"
And the idea that humans can't hear over 20khz is like "humans taste 'sweet' on the tip of the tongue, and 'bitter' on the sides"
As we get older the hairs in out ears break or whatever and our perception decreases, but I could hear the fly backs in my old monitors, I used to be able to see the flicker in 3khz pwm LEDs, and my induction hob drives my kids crazy but it's merely midly annoying to me.
Get a real soundcard and some young people and play square(pwm) and sine tones starting at 16khz and find out where they can't hear it anymore. I find studio monitors with tweeters that are not paper are the best.
The extra headroom can indeed be useful for some kinds of processing, but you can safely discard it for actual listening.
This seems to be mixing up two things; proper interpolation and dithering.
If you have limited bit depth (in practice, 16 bits or worse), you should pretty much always dither, ideally also noise shape. This is independent of the interpolation you're using; having a rational relationship between the original and downsampled signal makes some of the implementation a bit easier, but even for something like 48000 -> 24000, you'll end up with effectively a float signal that you need to convert to your chosen bit depth somehow, and that should be done better than just truncating/rounding.
And even for interpolating between two prime rates, or even variable-rate interpolation, you can and should get great interpolation (typically by picking out polyphase filtering coefficients from a windowed sinc of some sort).
Are they the exact same volume? We perceive things slightly louder as higher quality.
Is it a double blind test, ie an ABX test?
Are the bit depths the same? Many 96khz sampled files use 24 bits per sample, whereas 48khz usually uses 16 bits per sample.
but you do need phile-enough gears(minus the gilded pebbles hot glued onto circuit breakers)
The big issue with analogue landline phone calls is the audio bandwidth is so limited. It's not the full frequency spectrum, most of it it cut off.
https://en.wikipedia.org/wiki/Comparison_of_audio_coding_for...
https://en.wikipedia.org/wiki/Opus_(audio_format)#Quality_co...
EDIT: I do agree that lossless (or at least high bitrate modern lossy, like 256k Opus which is basically transparent) should be available in many more situations though.