I think you're right -- we mostly only talk to 1 person at a time.
Where it might be useful is automated live captioning (already a feature in Microsoft Teams), especially in cultures that favor a "high involvement" conversation style where people interrupt and talk over each other (vs a "high consideration" style where everyone waits their turn). This tech might help improve transcript quality.
In the article, under the subheading Why It Matters, the authors say this:
[Quote]
The ability to separate a single voice from conversations across many people can improve and enhance communication across a wide range of applications that we use in our daily lives, like voice messaging, assistants, and video tools, as well as AR/VR innovations. It can also improve audio quality for people with hearing aids, so it’s easier to hear others clearly in crowded and noisy environments such as parties, restaurants, or large video calls.
Beyond its separating different voices, our novel system can also be applied to separate other types of speech signals from a mixture of sounds such as background noise. Our work can also be applied to music recordings, improving our previous work on separating different musical instruments from a single audio file. As a next step, we’ll work on improving the generative properties of the model until it achieves high performance in real-world conditions.