It's interesting cause I have a recording of human voices plus a background TV show that was too loud; I've looked around for something that would be able to separate the two but I haven't found a straightforward solution.
For example if you Google then FASST is one of the ones that come up, but it's a whole framework and in order to use it you'd have to learn the research yourself; much of these software is not geared for end users.