I've thought about this problem before (but I have so many projects going on that I axed it).
Hypothetically, this is how I would approach it.
I would start by forking Open-Unmix or another 4-stem model (Demucs, MDXNet). The code is oriented towards the 4 sources (vocals/drums/bass/other), which is dominant because the major datasets for training use these stems.
The Open-Unmix training code goes like:
```
x, y1, y2, y3, y4 = load_training_data()
y1_est, y2_est, y3_est, y4_est = unmix(x)
loss = loss([y1, y1_est, y2, y2_est, y3, y3_est, y4, y4_est])
```
In your case it may be simpler like `noisy mix = clean mix + background noise`.
In that way, I don't even think you're constrained to using a dataset that has 4 stems available (which is a rare quality in a dataset).
Instead, you need a way to acquire or generate screaming and other concert noises.
Then, the new training code could look like:
```
x = load_training_data()
n_samples = x.shape[-1] # get length of music waveform
noise = generate_screams(n_samples)
x_noisy = x + noise
x_est = unmix(x_noisy)
loss = loss(x, x_est)
```
Well, anyway. That's my naive first idea of how I would approach it. But, this relies on having clean background noise/screams _without any music in it_. And I'm not aware of such datasets.