I've had an idea to use propellerhead recycle to chop the output cloned voice into syllables, and then "play" the chopped parts in rhythm, through autotune.
The issue is you get Eifel 65 sounding autotune if your base vocals are monotonic or way off key. The only way I can think of fixing this is to use something like audacity's pitch changer that doesn't affect the speed of the sample - rough the lyrical tones in with audacity/recycle, then autotune it where it needs to go.
I'd like to say I'm too busy to get this workflow going, but mostly I'm lazy and someone else will do it first - and better - I can't improve the AI cloning software.
I've had a new version of this in the works for a few months. It'll launch maybe this weekend? It supports user-trained data sets, which should be pretty neat.
Your approach sounds kind of like unit selection or vocaloid. Do it!! :)
I use(d) "Real time voice cloning" available on github. I have used your site before, i think that's actually the reason i wanted to find software to do it myself originally!
I ended up using mine to do the announcements on all of the ham repeaters in Central Louisiana. I'm not sure how many are still active, but i had a couple of podcasters and two presidents doing the announcements for a while.[2]
Nice work, wrapping it all in a website/service. I'll bookmark it!
[0] https://soundcloud.com/djoutcold/gottfried1-1
I love your use case for ham radio. That's brilliant!
And thanks. I'm glad there's other people messing around with voice cloning, still.
It's coming soon.
I'm working on mocap, photogrammetry, and compositing on Twitch, and I'm leveraging vocodes for subscriber interaction in the stream.
Are either of you working in this domain? Want to chat?
Check out Ableton Live’s “Simpler”
It does contain lyrics, but the singing is not AI generated.