Bark – Text-prompted generative audio model
github.com
github.com
Text to speech was a natural playground for us to share with the community and get some feedback. Given that this model is a full GPT model, the text input is merely a guidance and the model can technically create any audio from scratch even without input text, aka hallucinations or audio continuation.
When used as a TTS model, it’s very different from the awesome high quality TTS models already available. It produces a wider range of audio – that could be a high quality studio recording of an actor or the same text leading to two people shouting in an argument at a noisy bar. Excited to see what the community can build and what we can learn for future products.
Please let us know with any feedback, or if you’re interested in working on this: bark@suno.ai
For reference, look at how the societal damage of social networks has been handled: too little, too late. Same goes for RentTech.
But, I don't know the solution. The common computer has become so powerful that we cannot simply rely on inaccessible materials to prevent the danger of overly-powerful tech spreading too fast, as we do with bioweapons or traditional WMDs.
Fight tech with more tech, I suppose.
We want to go back to New York in the 90s? Petty theft everywhere?
With great power comes .... ? Profit?
Of course, you wouldn't encode real-time navigation data, but a small block of identifying text. Either way, though, someone without a copy of the spreading code isn't going to notice it or decode it. Given enough redundancy in both the time and frequency domains, removing it wouldn't be easy either.
PyTorch themselves used nanoGPT training as demo for this: https://pytorch.org/blog/accelerating-large-language-models/
Reference:
1. https://www.laika.berlin/en/blog/new-client-barkgpt-ai-dog-b...
Choosing to ignore your dog won't change because some magical AI can now translate it to -Im fine, Im only barking because you're an awesome being, keep your subscription humaaan-
Bark-GPT's VC pitch: "Replika for real dogs"
The other side of that problem is an opportunity. That's why the same model can also generate music, background noise and sound effects. And it's just because the prompt specifies those things explicitly. The input is truly semantic, so the output is rich and reflects that context. Is your input text sounds like it came from a speech, then there's a high chance your output audio will sound like a megaphone in a public space with crowd reactions and maybe even applause.
I'm not being negative -- some of the samples are really neat on their page -- and I know there is some idiosyncrasy of my setup that is causing issues, though it is a pretty typical conda + pytorch with CUDA 11.8.
Playing with the text and waveform temp from their defaults 0.7 is yielding some semi-decent results, but it feels essentially random.
https://www.youtube.com/watch?v=XSoOrlh5i1k
Someone needs to hook create the plumbing to capture speech to text, feed it to a GPT script that has been told how to reply to such call center calls, then send that back through a TTS generator like this one.
To overcome any latency issues, it could build in a ploy to buy time like the old script did, eg, make the robo-answerer sound like a somewhat addled old man who has to think before each reply, perhaps prefixing responses with "hmm, ahh, ..." to buy time to generate the response.
CYA aka https://en.wikipedia.org/wiki/Cover_your_ass
also it still requires tons of money to run, so it's likely only businesses will do it
Just like a private individual can own lock smith tools and play around with locks... don't go basing your business of supplying them wholesale to the general public worldwide for free.
That surprises me. Would you be able to point some out to me?
From what I've used of local LLMs they seem nearly unusable.
All the open source models I've seen so far have this weird kind of neural fuzziness to them. I don't know what Eleven does better, but there's definitely a big difference.
I guess, the "open" part of it is mostly for marketing.
Isn't this open source and can be easily removed or am I missing something?
Edit: looking at https://www.reddit.com/r/singularity/comments/12udgzh/bark_t... now.
It seems to be easily reproducible if I specify a non-existing speaker?
audio_array = generate_audio(text_prompt, 'en_speaker_3')
https://suno-ai.notion.site/Bark-Examples-5edae8b02a604b54a4...
>Bark has the capability to fully clone voices - including tone, pitch, emotion and prosody. The model also attempts to preserve music, ambient noise, etc. from input audio. However, to mitigate misuse of this technology, we limit the audio history prompts to a limited set of Suno-provided, fully synthetic options to choose from for each language.
It's not immediately clear how the audio history prompts are created.
from scipy.io.wavfile import write as write_wav
write_wav("c:\\wherever\\whatever.wav", SAMPLE_RATE, audio_array)
If you want to play it right away, you can use: from playsound import playsound
playsound("c:\\wherever\\whatever.wav")
What I wasn't able to do was make torchaudio play it directly, without creating an intermediate .wav file. Apparently there is a live-playback backend on the Mac, but not on on any other OS.