SoundStorm: Efficient Parallel Audio Generation
google-research.github.io
google-research.github.io
Then mocap, mapping digital faces on real actors which was first mind-blowing to see in Pirates of the Caribbean, then the apes in one of the Planet of the Apes movies... So much in the CGI industry has already reached a point where the hardest problems seem to have been solved.
When I now clicked play on the first Synthesized Dialoge from Dialogue Synthesis "Where did you go last summer? | I went to Greece, it was amazing.", I was blown away. It's as if we've now reached one of those milestones where a problem appears to be fixed or cracked. Machines will be able to really sound like humans, indistinguishable from them.
10-5 years ago, if you wanted to deal with TTS, the best option you had was to let your Android phone render a TTS into an audio file, because everything else sounded really bad. Specially Open Source stuff sounded absolutely horrible.
So how long will it be until we will be able to download something of this quality onto a future-gen Raspberry Pi which can do some AI processing, where we make an HTTP call and it starts speaking through the audio out in a perfect voice without relying on the cloud? 5 years?
If you think it is because there was no alternative, then you will see no barriers to the adoption of the sort of system you imagine.
But if you think it is because all human art is a form (however weird, tangled and opaque) of story telling, and for it to work, the person experiencing the art needs to be believe that there must be a story to be told, then your imagined system is not that interesting.
However, we are seeing an overwhelming propensity even today for people (even people who ought to know better) to ascribe intentionality and even emotion to computer systems that demonstrably do not have it (and as far back as Eliza in the 1970s). It seems likely to me that most people will rapidly come to believe in a sufficiently intentional and emtional backstory to your putative wunder-singer that their belief will satisfy their desire for "a person behind the art".
Both human and software-generated singing exist today, and either can be boring nonsense or speak to you and leave an emotional impact.
(That said, if said machine is actually a derivative work produced from vocal samples without appropriate licensing, that’d make it dubious legally or morally. To my knowledge, preexisting software-generated singing like Vocaloid does not suffer from this problem.)
The production of art that people want to experience probably requires a self, but may only require the belief that a self was involved.
The experience of art requires a self, as do all experiences.
that there are things made which are art does not imply that anything made is art.
that things can be made does not imply that everything is made.
You have seen my definition. By definition an act of self-expression cannot occur without a self.
Without clarifying what makes something art, your claim that art can be something produced without a self (i.e., is not self-expression) demotes art to the level of “any random thing”.
“You are a castrato who sailed with Black Beard and survived the holocaust with a degree in Klingon studies from Oxbridge…”
Not as an app exactly, but you should check out Holly Herndon and Mat Dryhurt’s suite of tools called “Holly Plus”:
I’m pretty sure you can access their model somehow and even train your own voice using their “spawning” approach.
She did an awesome TED talk demonstrating this:
https://www.ted.com/talks/holly_herndon_what_if_you_could_si...
Here’s a cool example, using Dolly Parton’s song “Jolene”:
https://www.youtube.com/watch?v=kPAEMUzDxuo
I don’t think it’s quite at the level of consumer use yet, but I know they’re working on it. Definitely check it out.
Virtual singers are already pretty good anyways, I feel we've already passed the point of diminishing returns.
And even if you look at virtual singers, the ones which are popular are the ones from ten years ago, not the newer ones with more realistic voices.
How far back are you counting as these days? I'd the entire recording industry has always been about attaching an image to songs.
I saw somewhere that this "girl" fills concerts.
- art is the process of creating
- art is the outcome of creating
I know which one I'd pick.
There are rights issues if the result is it replaces a particular singer, if you made it so that Sneaker Pimps can fire Kelli but still have her voice on subsequent songs that's a problem. But suppose you're a bedroom musician, and you realise you've got a piece that really wants somebody with a different voice than yours to make it work - you can pay someone, but technology like this offers a cheaper, easier option.
5 years? It's probably possible roughly whenever the larger Whisper models can run on it. Probably the next Raspberry Pi, running quantized or optimized versions of some audio model.
It may be almost possible right now if you tried really realy hard, and you used a small model fine-tuned on a single voice, instead of something larger and more general purpose that can do any voice. I think whisper-tiny works on a Pi on real time, right? And that's not leveraging the GPU on the Pi. (https://github.com/ggerganov/whisper.cpp/discussions/166)
Edit: looks like medium is 30x slower on the Pi than tiny model, so I may have been overly optimistic. I didn't realize Whisper tiny was that much faster than medium.
This method works pretty well with Tortoise, letting you use the super fast Tortoise quality settings but get quality similar to the larger models. Fine-tuning the whole thing on just one voice removes a lot of the cool capabilities of course. With Tortoise, that would still be way too slow for a Pi but potentially that same strategy could work with faster models like SoundStorm.
In terms of quality there's still a lot of room to go with long term coherence, like long audio segments. When a real person reads an audiobook the words at the top the page have a pretty big impact on how many words at the bottom the page are read. And there can be some impact at any distance, page 10 to page 300. When you try audiobooks on super high end TTS models and listen carefully you really notice the mismatch. It's like the reader recorded the paragraphs out of order, or a video game voice lines where you can tell the actors recorded all the lines separately, and were not reacting to each other's performance.
You can bump the context windows, a minute, two minutes. That's gonna get you closer and probably good enough for some books. In the short term a human could simply adjust all the all the audio samples and manually tweak things to sound correct. So this will enable fan-created audiobooks where they take the time to get it right. But for fully automated books the mismatch drives me nuts. The performance is just soooo close for certain segments that when you get a tonal mismatch it hurts.
These days, though...every new technique developed to simulate and replicate human creativity and behaviors builds on a constant sense of unease.
If I view or read, do I have the right to know if it's generated?
Bard's TTS is alright but it's clearly behind.
On that note, Bing's English/Korean TTS is really good. I also didn't realize Microsoft uses the best offerings for free TTS on edge so it blows google's default tts voices away.
Some of Azure's voices are better than others, and the TTS web app has a few minor bugs, but overall I was really pleased with the whole experience.
https://cloud.google.com/text-to-speech/docs/wavenet#studio_...
Just that
1. It's behind their current sota research
2. You can only use those voices extensively by paying for it. Microsoft offers their best stuff on edge for free. So for reading aloud a pdf or web page, microsoft is far better.
Without going into too much details, imo they’re not really usable right now for TTS use cases.
also Google Studio Voices is excellent. Definitely better than Microsoft's best, albeit very limited voices.
This sounds really interesting - can you share a bit more? I'm behind in this space, my parser got all jammed up, something like: "Microsoft uses [the best offerings for free TTS](as in FOSS libraries, or free as in beer SaaS?) [on edge](Edge browser, or on the edge as in client's computer?)(Is the implication that all TTS on the client's computer blows Google's default TTS voices away?)"
This is not the case with Google
sigh. Google used to release _some_ models. Guess the fun early days are coming to an end.
Even the demo website is on Github Pages instead of a Google domain/blog.
I think their non-ads businesses alone would be the 6th largest US tech company by revenue. (Amazon, Apple, Microsoft, the ads business of Alphabet, Meta. Am I forgetting something?)
https://abc.xyz/assets/investor/static/pdf/2023Q1_alphabet_e...
Maybe a third or a bit more of Bark outputs are a dialog person talking to themselves -- and it often misses a voice change. But the pipe characters do reliably produce audio that sounds like a dialog in the performance style.
https://twitter.com/jonathanfly/status/1675987073893904386
Is there some text-audio data somewhere in the training data that uses | for voice changes?
Amusingly, Bark tends to render the SoundStorm prompts sarcastically. Not sure if that's a difference in style in the models, or just Google cherry picking the more straightforward line readings as the featured samples.
I wonder how ML trained on the tone transitions to a sponsored segment dripping with secret shame... would infect general speech.
Their current marketplace interface seems inadequate for this. Instead of contacting a human and then wait for them to finish the work, buyers will want to get results right away.
Therefore they will have to change their platform to work like an app store. Where the sellers connect their services and buyers can use these services.
While those people are now looking for jobs elsewhere.
Instead, make an effort to start a conversation with your fellow travelers, or graciously respond to such effort from them. Apologize if you already do.
Likewise in many countries, having someone at the gas station to fill in the tank is still a job, and most likely even with EV they would be the ones taking care of the charging.
As do everyone that buys Nike.
Of course that wasn't a net loss, it was part a larger economic transformation that created more higher paying jobs.
Specially the people living from paycheck to paycheck, without any kind of healthcare support.
We will not run out of productive things to do with our time. Labor force participation has stayed in 60-70% despite centuries of automation.
At least until they get clever enough to start a transformers line factory.
There was never a fixed number of jobs, there's a fixed number of workers.
Isn’t this the ‘problem’ that AI is trying to solve?
Their users are already using AI to do the work that they are supposed to do. i think that's fine
Having a tool at my fingertips where I can feed it some of the actor's previous lines and be able to belch out something to fill the gaps with set parameters and be able to move along in the project without all the logistics would be heaven.
It would however, kill an entire field of expertise. It would also devalue the actor as well. Though it's already happening. There are already programs on the market that replace voice actors all together and are being used in the video game space.
For the work I do, I can see the benefits it could bring. But I am also fully aware it probably will be heavily abused.
The web does not have good discovery and reputation management and also does not provide a unified interface. That is why market places like Booking.com, Amazon, Spotify etc have become so big.
Am i missing something?
It just makes me sad that I cannot open this page on Safari. It will not play a single audio, yet Chrome plays it fine. So here we are, able to generate audio, video, code, do amazing things with AI, but a simple website that has text and audio is not working on the most popular laptop out there.
I understand that some applications might prefer "other browsers", but a simple page that plays audio snippets was really a disappointment.
But think open world games. GTA VII for example where all NPCs have their dialogs auto generated in real time but also converted to audio in real time.
That's going to be a world which would be a lot more spontaneous with lot less effort.
Right now, If memory serves me right, GTA V dialogs alone are 5000 pages or more, hand written.
This is applicable to LLMs as well. You can get it to write plausible BS but if you really want a rooted in reality, well articulated write up about something, a human has to be taken onboard.
This equally extends to voice over. If you really want expressive and creative control to put some outstanding rendering of something, AI isn't going to cut it.
I welcome any tech that makes indies more competitive.
You are saying it’s not economical to use tech to speech to support blind people. I’m saying the benefits are huge for older population. It isn’t just for fraudsters or spammers as you claim.
None of the objections voiced to my original comment have even attempted to engage with this economic industrial reality.
Imagine replacing crappy phone menus with polite virtual assistants that actually understand what you're saying.
Imagine an AI language tutor that speaks every language in the world fluently. Or a universal speech-to-speech translator.
And that's just off the top of my head. Clever people will come up with a lot more uses, I'm sure.
It will be interesting to see what happens when that pressure is applied...