HNHacker News
TopNewBestAskShowJobs

sammyyyyyyy

173 karma · joined October 11, 2024

submissionscomments
sammyyyyyyy··on [dead]
Today, I'm releasing a new version of my side project: SoproTTS

A 135M parameter TTS model trained for ~$100 on 1 GPU, running ~20× real-time on a base MacBook M3 CPU.

v1.5 highlights (on CPU):

• 250 ms TTFA streaming latency

• 0.05 RTF (~20× real-time)

• Zero-shot voice cloning

• Smaller, faster, more stable

Still not perfect (OOD voices can be tricky, and there are still some artifacts), but a decent upgrade.

sammyyyyyyy··on I trained a 135M TTS model for ~$100, runs 20× real-time on CPU
Nice
sammyyyyyyy··on I trained a 135M TTS model for ~$100, runs 20× real-time on CPU
Today, I'm releasing a new version of my side project: SoproTTS

A 135M parameter TTS model trained for ~$100 on 1 GPU, running ~20× real-time on a base MacBook M3 CPU.

v1.5 highlights (on CPU):

• 250 ms TTFA streaming latency

• 0.05 RTF (~20× real-time)

• Zero-shot voice cloning

• Smaller, faster, more stable

Still not perfect (OOD voices can be tricky, and there are still some artifacts), but a decent upgrade.

Repo: https://github.com/samuel-vitorino/sopro

sammyyyyyyy··on Sopro v1.5: A 135M TTS model trained for ~$100, runs 20× real-time on CPU
Today, I'm releasing a new version of my side project: SoproTTS

A 135M parameter TTS model trained for ~$100 on 1 GPU, running ~20× real-time on a base MacBook M3 CPU.

v1.5 highlights (on CPU):

• 250 ms TTFA streaming latency • 0.05 RTF (~20× real-time) • Zero-shot voice cloning • Smaller, faster, more stable

Still not perfect (OOD voices can be tricky, and there are still some artifacts), but a decent upgrade.

Repo: https://github.com/samuel-vitorino/sopro

sammyyyyyyy··on Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
You should try it! I wouldn’t say it’s the best, far from that. But also wouldn’t say it’s terrible. If you have a 5090, then yes, you can run much more powerful models in real time. Chatterbox is a great model though
sammyyyyyyy··on Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
Also, I didn’t want to use known voices as the example, so I ended up using generic ones from the datasets
sammyyyyyyy··on Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
I should have posted the reference audio used with the examples. Honestly it doesn’t sound so different from them. Voice cloning can be from a cartoon too, doesn’t have to be from a human being
sammyyyyyyy··on Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
I didn’t specially cherry pick those examples. You can try it anyway for yourself. But thanks for the feedback anyway
sammyyyyyyy··on Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
As I said, some reference voices can lead to bad voice quality. But if it sounds that bad, it’s probably not it. Would love to dig into it if you want
sammyyyyyyy··on Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
Yeah sure. The training was about ~250 dollars, which is quite low by today’s standards. And I spent a bit more on ablations and research
sammyyyyyyy··on Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
No, it doesn’t.
sammyyyyyyy··on Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
Yes, you are right. However, there are many upsides to this kind of technology. For example, it can restore the voices of people who were affected by numerous diseases
sammyyyyyyy··on Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
Obrigado! Quando (e se fizeres isso) manda pm!
sammyyyyyyy··on Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
Yeah, we are not quite there, but I’m sure we are not far either
sammyyyyyyy··on Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
This is my side “hobby”. And compute is quite expensive. But if the community’s responsive is good, I will definitely think about it! Btw, chatterbox is a great model and inspiration
sammyyyyyyy··on Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
Cool! Yeah the voice quality really depends on the reference audio. Also mess with the parameters. All the feedback is welcome
sammyyyyyyy··on Sopro TTS: A 169M model with zero-shot voice cloning that runs on the CPU
Thanks! Yeah I kinda postponed publishing it until it was a bit better, but as a perfectionist, it would have never been published