But we also rarely need realtime solutions for this. Casually tracking gross features from a mixed audio source, like a big crowd, is easy, but we almost never need to be actively listening and disambiguating between several deliberate audio signals all at once. Human communication just isn’t setup that way, though I’m sure there are niche exceptions.
Note here that video conferencing is emphatically not an example of an exception. You still need to have synchronous order of speaking, not because technology lacks the ability to separate the streams for the listener, but because the listener cannot pay attention to more than one stream at a time.
To me this technology seems almost exclusively useful for surveillance and offline audio analysis or audio synthesis / mixing.
Still possibly valuable, but definitely not in any fundamental way that would significantly change or augment realtime verbal communication.