From what I can tell it seems the RNN learned when the speaker was talking. It then just make sharp cuts to blank out when the speaker is not talking. It does not appear to have learned how to extract just the frequencies of the speaker but rather just when a speaker is speaking.
I feel this could be taken a step more in such that when the speaker and overlapping loud sound happens at the same time it is able to extract just the speakers voice.
Now obviously this is easier said than done.