HNHacker News
TopNewBestAskShowJobs

piotr11

22 karma · joined February 18, 2022

submissionscomments
piotr11··on This Voice Doesn't Exist – Generative Voice AI
They will be multi-lang, the tech scales to any language and we are working to add more (it is relatively easy). Here is the demo in Polish TTS: https://www.youtube.com/watch?v=ra8xFG3keSs
piotr11··on This Voice Doesn't Exist – Generative Voice AI
Exactly (ElevenLabs dev here)! This is actually out mission - make all content available in any language and voice. Dubbing is where we are going!
piotr11··on This Voice Doesn't Exist – Generative Voice AI
Thanks (ElevenLabs dev here), we are constantly working on improving our model, we do out own research and train it completely from scratch.

We do support Polish already and the quality is actually better IMO than English as we use a newer generation model: https://www.youtube.com/watch?v=ra8xFG3keSs Some people think it is fake and we hired a real voice actor to read.

piotr11··on This Voice Doesn't Exist – Generative Voice AI
Thanks! ElevenLabs dev here - these are generated 6x faster than real-time, with latency of <1s. No corrections required.

We are working on long-form speech synthesis too, needless to say, the audio reading the article has been also synthesized by a voice that does not exist.

piotr11··on This Voice Doesn't Exist – Generative Voice AI
Hey - developers behind ElevenLabs here. Thank you so much for the constructive and positive feedback - we’re taking it onboard!

We’re currently focused on researching and deploying a different way for speech synthesis that can generate nuanced intonation and emotions by understanding text and taking context into account. Additionally, we provide creators with a way to clone their own voice based on very short samples. With the published blog post, we are now deploying a way to help them design entirely new ones!

Anyone will be able to generate that level of quality just with a copy-paste. We are planning to open up Beta later this month. Our goal is to let you convert any written content into high-quality, compelling audio.

To address a few questions that frequently came up:

- Latency for our streaming TTS is <1s with quality results available above, which is the usual problem with existing good TTS models (like tortoise-tts)

- We can clone voices instantly, based just on 5s of speech, without training required

- We are working on adding SSML-like support for better control; speed controls will be coming as part of that too

- API is directly available as part of Beta; we are preparing the infrastructure to scale easily for the release!

We are hiring researchers, frontend and full-stack developers! If you are interested, send over your GitHub account and short message to founders[at]elevenlabs.io.