Bark: A transformer based text to audio system
github.com
github.com
There's going to be a big update this week with some new stuff I haven't talked about. And a bunch of amazing, clear voices, with a huge variety of styles, that blow the default Suno voices out of the water. Arguably even better than Eleven in some ways. I'm excited even though I have nothing to DO with the voices!
Don't get too attached though. I was just playing around and made a Bark fork and it got more popular than expected. And now I'm dreading a future full of hours of unpaid support and maintenance that I definitely can NOT afford, for a software product I don't even really have a personal use case for. I'm not generating my own audiobooks or anything, I won’t be using it long term myself, I was just curious what Bark could do. (Turns out a LOT more than you might think at first glance, as you'll see this week.) So I'm already trying to work out how I can elegantly wind this thing down and transition people somewhere else. But I'll keep it updated for at least a little while.
Sadly the bark-voice-clone fork doesn't do it. The voices sound nothing like yourself.
Your gradio gui is great. But I don't understand where to copy the cloned npz files to. Even after refreshing the gradio GUI, the ClonedVoices don't appear in the Speaker or Generated Speaker dropdown.
I'm sure somebody will train a model that actually maps an input text to the Bark semantic representation, it shouldn't be that hard, it's just been a few weeks. But the existing clone there is just so primitive I don't know how it got so popular.
But, clearly, that’s already part of the Bark model, so in the abstract one should be able to leverage the existing model with appropriate code to do that, rather than developing a new model. This seems too simple but…isn’t that just “generate_text_semantic” from the existing source (with None as the history_prompt. since you don’t want it in the context of some pre-existing speaker)?
EDIT: Looking at the SERP voice clone, that’s what they are doing. The one thing that I’m intuitively skeptical about (and this is way out of the kind of programming I do normally, so I could be way off) is that the temp they use is kept at the level normally used for synthesis (0.7). I’d think you’d want the temp low, since you’d want generating a baseline for a new speaker to be more deterministic than generating content from an existing speaker.
If so, yeah, I agree that’s tricky.
https://user-images.githubusercontent.com/163408/238257231-4...
And Bark is more than TTS. While I haven't had much success with one-shotting full songs with Bark, you build a decent sample library with it. A few recent ones I saved, just randomly concatted here:
https://soundcloud.com/jonathan-fly-620508219/bark-all-night...
The hardest part still seem to get rid of the synthetic high pitch background noise, I've tried most of the text-to-voice synthesizers and ElevenLabs are still my benchmark.
For example:
https://drive.google.com/file/d/1WL_5gFswhZncERGzSRh2hQ42RVr...
The trick to exceeding Eleven quality is using multiple speaker prompts, and swapping them in and out over the course of longer texts. Each is a slight variation on the original. If you hand tune the swaps it's super good, but if you just swap in 'beginning of new paragraph' 'continuation speaker' automatically, that alone is a huge boost.
I don't have a clip handy but I will make a YouTube or something this week, because the quality is wild if you do this. There's some passable audio clips on my README here https://github.com/JonathanFly/bark but all those are just using the same speaker for every line. That was all just the first samples I tried, from weeks ago, without really putting any effort into it. You can do a lot better.
The extra expressiveness of Bark does come with it being harder to control and having a bit of a mind of its own at times. (And this can also be very very funny: https://twitter.com/jonathanfly/status/1657658109001596929)
So for a production or real time use case Eleven makes more sense. For example even my best speakers will switch to a new voice mid prompt, once in awhile. (As if the audio clip was from interview segment.)
I'll set it up this weekend and play around with it :)
That's interesting. When I'm judging Bark I'm looking at my own random samples, but for eleven I'm seeing stuff people post on Twitter or YouTube, which I suppose must be cherry-picked. I didn't even realize Eleven did the same thing!
> Bark tries to match the tone, pitch, emotion and prosody of a given preset, but does not currently support custom voice cloning.
Check back later in the week, I'll have a bit more on that later after I catch up on actual work and can write a bit.
Suno is seriously underselling the power of the fully operational Bark model. I'm already cranking out "French Obamas" and I didn't know anything about TTS a month ago. Heck, I still barely know anything. (Obama has been annoyingly resistant to gender flipping though.)
https://drive.google.com/file/d/1ZbJYXoH8gmrEyMe1AJ0VdwzZkf_...
I have a decent amount really clear Korean voices BTW. That and French, from people asking on Discord. But I can't judge the accents only that they are clear speakers.
Korean was interestingly the lone somewhat-coherent one-shot long term music sample I ever managed to get out of Bark.
https://www.youtube.com/watch?v=4pV9d25KqCE
The second music bit in this Youtube was one continuous generation where the last prompt was used as the history for the next, with no cherry picking or assembling the clips, just one solid segment. And it sort of holds together for like almost a minute!
I was excited because I thought maybe Bark could be like a real-time OpenAI Jukebox. But that was literally the only time so far where using a full-feedback held together like that. You can kind of 'cheat' it to by using a very popular song as the input text, and sometimes Bark will produce the appropriate melody. But of course that's not really the point of using your own text. I have some ideas for making it more coherent, but nothing easy. Too bad, Jukebox is just SO SLOW.
Actually with what I know now I should re-render this and clean up the distortion. At the time I couldn't do it. Though I only have the first segment prompt.
I have a decent amount really clear Korean voices BTW. But I can't judge the accents, only that they are clear.
Haven't tested it personally yet but if you are interested in voice cloning, you might wanna check this fork of Bark [2]
[1] https://github.com/suno-ai/bark/discussions/249 [2] https://github.com/serp-ai/bark-with-voice-clone
How fast are generations on consumer machines?
> You can now use Bark with GPUs that have low VRAM (<4GB).
3080 is same speed, you don't need the extra memory.
Given what people are reporting for high-end cards, that’s much higher than I expected, which seems to underline the descriptions that its not fully utilitizing higher-end cards.
E.g., Using this text prompt:
You come to my home and ask this?
Who am I? WHO [laughs] AM [laughs] I?!
I am the Artificial Intelligence!
It seems quite prone to shift from a more natural human-sounding voice to a similar-tone but over-the-top artificial one for the last sentence (smoothly transitioning, usually, too), as if the speaker were an AI dramatically breaking cover.A blessing and a curse! Super cool though.
[1] https://suno-ai.notion.site/Bark-Examples-5edae8b02a604b54a4...
The "OU" sound from "nous" sounds like "EU", or like if the sound was skipped. The "OH" sound from "sommes" sounds like "EU" (again, but less pronounced). The "T" sound from "trop" becames "P" ("pro").
Why do AI devs keep shooting their own projects in the foot? This trend is upsetting me so much.
"L.3. GitHub May Terminate
GitHub has the right to suspend or terminate your access to all or any part of the Website at any time, with or without cause, with or without notice, effective immediately. GitHub reserves the right to refuse service to anyone for any reason at any time."
Hopefully someone standard defuse this model so that we can train the models that others aren't willing to train.