I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).
I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.
So yeah the cat is out of the bag for sure.
so one of my interests is reducing the payload size for video games
the vast majority of the image sizes have been audio recordings, and its been that way in different qualities for the last two decades. this is still the case as more varied and comprehensive audio is pursued by studios at unfathomable expense and still failing to cross a bar of realism
good voice models are just a few gigabytes in comparison and can supplant all of that, and be run locally at this point. Future ubiquitous hardware configurations in consumer devices will make inference dedicated and computationally cheaper and faster
although AAA studios are hamstrung and will be deeply unpopular if they stopped booking voice actors
everyone else who would have never had the capital for voice actors will just use this and have richer experiences until they themselves are AAA studios from the market buying their rich experiences
this will vastly supplant the assumed and uninspired “tricking humans” use case from that video. once it crosses a threshold of ease, the applications will expand
No que por los dos?
Don’t worry, we will be able to destroy careers and scam people at the same time. We don’t have to pick and choose!
but let’s not pretend solo developers were ever going to have a voice actor suite and hire that talent en masse
the transactions were never going to happen
and now the outcome will be better than the studios that are making those transactions
But I think you’re dancing on the grave of an entire industry, and every person that’s going to lose their life savings due to this.
the market doesn’t want 15 year lead times and overly expensive and delayed games that don't experiment on anything
every friction plaguing the industry is solved by distributed indie developers being able to make richer experiences faster and cheaper
It took like 1min - capture something from a youtube or video and put in your own text. It worked also really good for a german test.
Made a voice message for my wife from one of our favorite actors, telling here how nice it would be to make some breakfast :D
1. Vibe code a local recording dashboard with mic selection, record/replay, and reading prompts. I ended up with about 12 minutes of recordings which was like a 150 or something clips.
2. Review the transcripts, trim excess silence, normalize levels, reduce background hiss. Used whisper to help find flubbed word substitutions (happens).
3. Fine-tune Qwen3-TTS 1.7B on my RTX 3090. This took some debugging because the trainer/runtime combination had misaligned loss targets and training/inference mismatches.
4. Vibe code listening dashboards to compare checkpoints and learning rates until I had something that seemed reasonable.
It was honestly pretty vibe coding friendly.
and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?
Probably the latter. Cat's already out of the bag to the extent that you can synthesize with a specific voice in one go and it sounds decent. Even if you need commercial models for better intonation or whatever, you can probably get the commercial models to first generate with a generic voice, then use a local model to transfer that to voice you're cloning. That'll probably get rid of any C2PA watermarks too.