The OS doesn't have more information about this than applications and it's not that obvious whether an application wants the OS to fuck around with the audio input it sees. Even in the applications where this might be the obvious default behavior, you're wrong - since most listeners don't use loudspeakers at all, and this is not a problem when they wear headphones. And detecting that (also, is the input a microphone at all?) is not straightforward.
Not all audio applications are phone calls.
the OP pointed out that this only works if he uses a browser monoculture
the OS does have more information than that, it can know what is being played by any/all apps, and what is being picked up by the mic
fwiw, you only need to know anything about outputs if you are doing AEC. Blind source separation doesn't have that problem and can just process the input stream.
Even if this is true, it's easy to imagine such functionality being exploited by malicious apps as a security and/or privacy concern, particularly if the user needs a screen reader.
It definitely makes sense for the operating system to provide this functionality.
Really what would be nice is if every audio i/o backend supported multiplex i/o streams and you could configure whether or not to cancel audio based on that set of streams but not all output (because multi output-device audio gets tricky).
I'm sure there are some niche cases, but in those cases, the application can specifically request that the OS turn off audio isolation.
That latency is within the tolerance that users are comfortable with for voice chat, and much less than video processing/transfer is introducing for video calls anyway, so it's a very obvious win there. Especially since those users are most interested in just picking out clear words using whatever random mic/speaker configuration happens to be most convenient.
But musicians, for instance, are much more interested in minimizing the delay between their voice or instrument being captured and returned through a monitor, and they generally choose a hardware arrangement that avoids the problem in the first place. And that's not really a niche use case.
Default on vs default off is really just an implementation detail of the API though, as you say.
If I'm recording a voice memo, or talking to an AI assistant, I would want this. Basically everything I can imagine doing with a PC microphone outside of (!) professional audio recording work.
That last case is important and we agree there needs to be a way to turn it off. I think defaults are really important though.
As you say, as long as either option is available, the only question is what the default should be.
Music player, browser, games, video player...
Audio is not app specific
The only application were this is true is audio were you want full control and low latency.
I find your take very weird.
Some do.
But you need to have a strong-handed OS team that's willing to push everybody towards their most modern and highly integrated interfaces and sunset their older interfaces.
Not everybody wants that in their OS. Some want operating systems that can be pieced together from myriad components maintained by radically different teams, some want to see their API's/interfaces preserved for decades of backwards compatibility, some want minimal features from their OS and maximum raw flexibility in user space, etc
Which Operating systems do this?
Frankly, I imagine its also available at the system level on Windows (and maybe Android and Linux) but probably only among applications that happen to be using certain audio frameworks/engines.
1. https://www.freedesktop.org/wiki/Software/PulseAudio/Documen...
Wait, what other audio paradigms are there?
https://learn.microsoft.com/en-us/windows-hardware/drivers/a...
Being able to get exclusive access/bypass the system via certain means (ASIO would be another) doesn't make it go away.
This is the way things usually work in the Free Software world. For example: need JPEG support? You'll probably end up linking to libjpeg or an equivalent. Most languages have a binding to the same library.
Is that part of the OS? I guess the answer depends on how you define OS. On a Free Software platform it's difficult to say when a given library is part of the OS and when it is not.
My experience is the opposite. When it's part of the OS, it's stable and you just say "you need OS version X or better" and it will just work. When it's a library, you eventually end up in dependency hell of deprecated libraries and differing versions (or worst case, the JavaScript ecosystem when the platform provides almost nothing and you get npm).
That's mac of course but in my experience Windows is much more trusting of what it gives applications access to so I suppose the same thing is available there.
https://developer.apple.com/documentation/avfaudio/avaudiose...