My first internship was at a company that was doing this over a decade ago, but before machine learning had proliferated to adjacent fields. It's easy to separate out stationary signals like ambient noise, which are easily isolated in the frequency domain. But it's a totally different thing to remove something like a baby's cry, shuffling of papers, or other sharp transients. Blindly separating without a large corpus of inputs + machine learning techniques seems impossible.
I remember as intern spending hours in front of spectrograms manually deleting noise so that the researchers could get clean targets. Let me tell you, I started being able to identify a lot of phonemes just by visually examining waveforms.
Eventually, the company did pursue some noise cancellation, but only as one part of their offering. I don't think they ever could get the holy grail of separating non-stationary noise.