Prime Voice AI: AI speech software
beta.elevenlabs.io
beta.elevenlabs.io
The actual sound quality of the output is impressive (clear treble, no weird artifacts between syllables, etc.), but I just don't understand the weird "edginess" of the speech.
But their model is really good when it comes to cloning voices from small audio samples. It was discovered by 4chan unfortunately [1]. I have only seen a few clips but all of them were racist, sexist or worse. Not appropriate to link them here, I guess. You can see an official sample on their YouTube channel [2]. However, the voices and conversations I've heard yesterday, other than being disgusting, were high quality, believable and full of emotions. The voices of Obi Wan Kenobi or Joe Biden sounded so genuine that it was creepy. I know there has been tools to deepfake voices for years now, but this is the first time I'm seeing one that sounds so authentic.
[1]: https://www.theverge.com/2023/1/31/23579289/ai-voice-clone-d...
However their pricing is completely wrong, should be cheaper and offer more.
Before you had free, 10k a month & $22 a month for 60k, now they added $5 a month for 30k and $22 a month for 100k.
Attribution for free is pointless, as with custom voices in theory it would be hard to detect and enforce anyways.
Once you've downloaded the Premium voices (e.g. Zoe) it's just a CLI, no API or hidden bells and whistles.
$ say -v 'Zoe (Premium)' "This is an example of the Zoe voice for my comment on Hacker News."
You'll have to download the voice ahead of time, but Zoe (public) and Maeve (internal) are both excellent voices.There are text extraction utilities for PDFs which reconstitute paragraphs and whatnot. Seems like an obvious thing to do. I suggested it, but didn't hear back.
I imagine that PDF munging skills aren't common and a solo developer doesn't have the bandwidth to be smart about so many different techs.
https://azure.microsoft.com/en-us/products/cognitive-service...
Plug warning - I've been using it to create narration for short stories with it for a while, and the output is better than I would have expected. Here's a recent example involving two characters talking - https://storiesby.ai/p/melancholy-musings-over-drinks
For narration it can work. But for dramatic character acting in animated films, the results make the characters sound like terrible actors. More granular control is needed over specific words, syllables, tone, emphasis and timing.
This would be a boon to those who have lost or will lose the ability to speak or speak well. Especially if it can be integrated into communication apps and ones cell phone.
The number of people who could use this is going up as the hpv+ head and neck cancer wave ramps up.
I can get almost the same quality using open source models. Plus I can fine-tune them to get custom voices. That means any company who needs TTS is cheaper off paying me once to build them a customized open source solution instead of forever paying this company per minute.
This Voice Doesn't Exist – Generative Voice AI - https://news.ycombinator.com/item?id=34361651 - Jan 2023 (260 comments)
I had to adjust the punctuation to get it to sound more natural, but it was surprisingly good. Way better than the AI voices I hear on YouTube videos.
i.e. sounds very natural and not like a robot from the 1980s and doesn't require a cloud service and can run on modest hardware