Show HN: VoiceClonr attempts to reconstruct human voices
voiceclonr.com
voiceclonr.com
Edit: Text to pronunciation is a whole other problem.
What hasn't been done well yet is extracting a model from existing uncontrolled voice samples. That's what this is trying to do. Once this works well, software clones of dead singers will be popular. The RIAA is going to hate this.
Aiming for lyrics is a much higher target than everyday text though, due to grammatical hints and the extra pitch and phrasing demands of lyrics. Your results might hit people harder on non-lyrical textual bodies.
Keep up the good work, I'd like to make something like this for musical instrument someday :)
P.s. have you ever heard of Douglas Hofstadters Letter Spirit project which synthesizes fonts from a subset? http://www.cogsci.indiana.edu/farg/mcgrawg/fonts.gif
Reminds me of this -- before Roger Ebert died, he tried to have his voice reconstructed by some company using audio from his TV show, etc., but alas, it was too difficult at the time, so he ended up using one of the Apple TTS voices instead.
https://www.youtube.com/watch?v=93jREDSWOYY#t=1m23s
Google "Cereproc" and "Roger Ebert" and you'll find that he was quite pleased with it.
--snip
In early 2010, Ebert and Chaz announced on the “Oprah Winfrey Show” that they’d enlisted a Scottish company called CereProc to create a computerized voice that more closely resembled Ebert’s own by using snippets of his TV work, DVD commentaries and the like, but that never fully materialized. Alex stayed with him until the end.
--snip
Alex being the Apple TTS voice I mentioned earlier.
Source: http://voices.suntimes.com/arts-entertainment/the-daily-sizz...
To illustrate: https://www.youtube.com/watch?v=v7-Gwg0rL0k#t=5m33s - this one's not particularly well done (and doesn't even do all those steps), but you see what I mean.
One of my biggest peeves right now is that voices cost a ton of money, few are readily available otherwise, and a lot of the new stuff is cloud-dependent. (Which is a big turn-off to me.)
What are you looking to do with this?
There's hundreds of episodes containing it, including remastered audio in the HD versions of TNG.
One suggestion: make the text-to-speech button bigger and centered (I missed it the first time).
Why even attempt to get things like imitating specific people's voices to work when your speech isn't even fluid and pronounced clearly to begin with?