It is amazing how the ear manages to distinguish all the sounds without distortion.
It is amazing how the ear manages to distinguish all the sounds without distortion.
That's like complaining about how bad the pig plays the violin. This is absolutely incredible. The complexity level for this problem is right off the scale and the software does a passable job of it. Given some time and more training data and a few more people working on it this has serious potential.
[1] https://pigcasso.org/wp-content/uploads/2018/12/7.jpg [2] Also, I am not affiliated with anything in particular that has been mentioned, or with the pig.
I have no idea how this tool splits them up at the implementation level but I imagine it tries to split it up based on frequencies and when it lifts out the voice, it's cutting out a ton of frequencies that would normally be in your voice so now it sounds very unnatural, blocky and metallic.
With studio quality headphones I can notice a massive difference for the worse between the original and separated vocal track. It reminds me of when I turned on a noise gate too high when recording audio for my courses. That noise gate clamps down on certain frequencies to help eliminate room noise but it also dampens or removes natural frequencies that occur on the low end of an average male voice. It gives that same very jagged sounding audio waveform.
For example, in the beginning, if you listen to the phrase "from the office ...", in original recording it sounds smooth, like a single phrase (and the voice is warm and pleasant), but in the separated track it sounds like it is synthesized from pieces that are not properly connected, like vocaloid songs. It sounds little harsher. Transitions sound unnatural. And the phrase "heya Tom" in the separated track is split into "heya" and "Tom" with some unnatural sound between (or maybe in the beginning of "Tom"). It is like transitions you can hear in vocaloid tracks. Or the kind of artifacts you get if you over-compress an MP3 file.
And "it's good to see you" part also sounds robotic.
Maybe it's losing some of harmonics, but in a different way for different syllables and they don't sound like a single phrase anymore.
The vocals don’t have any significant residual artefacts from drum hits or any residual bleed-through of instruments playing the same notes. It’s magical.
Edit: the source is an mp3, which removes audio frequencies based on perception/masking with other frequencies. It's perfectly normal that it is showing artifacts. A better source is needed.
Edit: I uploaded the original FLAC to my web server last night, before I decided that MP3 would be more convenient:
https://mwcampbell.us/tmp/spleeter-demo/jonathan-coulton-re-...
I basically want to run this over every Steve Gadd recording I have.
I think the next step would be to train a network that can un-robotify songs and then run it on this.
funny, the original recording seemed kind of robotic to me! maybe not robotic, but like it's been filtered somehow. but that might just be my not-so-great headphones