Voicebox: Generative AI model for speech that generalizes across tasks
ai.facebook.com
ai.facebook.com
Learning from the pushback on releasing LLaMA it seems. I wonder how hard this will be to replicate. (Trained on “60K hours of English audiobooks and 50K hours of multilingual audiobooks in 6 languages for the mono and multilingual setups”, this doesn’t sound intractable.)
> Model Transformer [Vaswani et al., 2017] with convolutional positional embedding [Baevski et al., 2020] and ALiBi self-attention bias [Press et al., 2021] are used for both the audio and the duration model. ALiBi bias for the flow step xt is set to 0. The audio model has 24 layers, 16 attention heads, 1024/4096 embedding/feed-forward network (FFN) dimension, 330M parameters. We add skip connections connecting symmetric layers (first layer to last layer, second layer to second-to-last layer, etc.) in the style of the UNet architecture. States are concatenated channel-wise and then combined using a linear layer. The duration model has 8 heads, 512/2048 embedding/FFN dimensions, with 8/10 layers for English/multilingual setup (28M/34M parameters in total). All models are trained in FP16.
It's very cool tech, but it's far from transparent. It has a very obvious "autotune" like sound to it that jumps right out. when they edited that one word it was obvious it had been edited.
Again, super cool tech, just not going to replace voice actors or anything.
Government ID doesn't help much with those -- it's actually the thing that is not strong enough.
There are alternatives to using the state for this, but they are difficult and fraught with UX issues. Perhaps a decentralised web of trust or some sort of blockchain based registrar of trust that can trace trust routes between mutually distrusting individuals.
Unless such a system is in place, international and strong before states start playing in this space, there isn't much chance of beating a state's approach to the problem.
Just look at https certificates. The current system involves browsers shipping configured to trust a whole bunch of entities I don't really trust, and there has been relatively little interest in trying to build a working decentralised approach to site security.
It would be nice if we could come to a thorough solution that actually does cover all bases, rather than all these companies trying to create their own digital ID services that just encourage us to instead do silly things like photograph your ID front and back.
I mean hell it's taken like 20 years for privacy by design to become an ISO standard? That sort of timeline is not something we can really tolerate as more and more people continue relying on online services and in turn wind up trusting horribly outdated techniques/general malaise about data.
To me it has a very obvious "Hindi is my native language" accent. I mean after literally the first sentence: "The research team at Meta is excited to share our work...". Ouch. The "our work": just ouch. I was wondering why it wasn't a native english speaker presenting the video when the video is precisely about generating speech.
The first seven seconds are particularly bad.
Don't get me wrong: I've got a lovely french accent when I speak english.
This has either been trained on too many audiobooks spoken by non-natives or they've used their own tech, where the "reference audio" given as input was from a non-native.
In any case something is seriously off.
At 1:59, the "Hi guys, thanks you for tuning in! Today we are going to show you..."... That is obviously an Hindi speaker speaking (it's an example of fixing a real voice by removing background sounds).
I think that the main voice of the video was done by the same person who did the example at 1:59. And I think that they used their example of using a "reference audio".
And that person ain't a native english speaker.
To compare: when the reference audio uses a proper english accent (the example with the "diverse ecosystem" at 0:52), then the output from the text-to-speech sounds native.
I think they just fucked the demo video and it may already be ready for prime time.
To me, this is a style choice for the demo. Not evidence that they "fucked" it up. Accents are common - everyone has one! It's nice to see the model can support your personal voice even if it's not completely neutral English.
There is no such thing as "neutral" English.
TLDR: "neutral English" is like "neutral water temperature" - it feels neither hot not cold because it matches ones body temperature. It's subjective, and terming it "temperatureless water" is even less accurate.
I'd put emphasis on "perceived" and "American" in that statement, and also note that this is limited to regional accents: General American is unambiguously American. Similar to General American, many countries have developed a "Newscaster" accent, e.g. Received Pronunciation for Britain, but it's not considered neutral as it is the "upper class" accent.
In every language I've known well enough to distinguish accents, I've realized newscasters adopt a distinct accent/cadence that's not commonly used. But I wouldn't call it "accentless" - it's just another accent that may/may not have evolved from a culturally dominant regional accent (or dominant figure from a specific region.)
When. Narrating. Videos. One. Tends. To. Speak. Differently.
Or, the more important case -- if I'm listening to audio-version-of-X, is it sufficiently human-like that I can forget that it's synthesized voice?
To me, yes.
Easy to tell if you're specifically listening for it, but to use an analogy one doesn't typically read novels and parse closely for grammar, does one? Your attention is elsewhere, on the content and plot.
Its more surprising nowadays if an article about AI doesn't have a "twist" that what you read/heard/saw was AI.
If this feature is standardized and built in it could be paired with a token like a yubikey which is on users keychain and authenticated even if they were using someone else's phone.
All these details can be completely opaque to the user who just needs to see a blue check by calls made with a device verifiably logged in as name@provider and a red exclamation beside calls made with no info at all.
Think fake kidnapping scams, industrial espionage, recent attacks involved in pretending to be in HR to redirect corp transfers to attacker.
Just a little doubt is liable to blow up such scams.
Remember the standard tech bait-and-switch: if this feature is built in, it'll not be good enough to function for purposes you describe, but it will be good enough to track you by your voice for advertising purposes.
Now, I just want to talk about my little weekend project... I spent a couple of hours scraping Royal Road and trying to get TTS working. Eventually, I settled on:
1. `wget --recursive` filtering only the chapters 2. A python script to strip extraneous html like advertisements and the headers. 3. Pipe into pandoc emitting plain text. 4. Copy it to my phone for TTS: https://f-droid.org/packages/com.danefinlay.ttsutil/
I really wanted to use all local tools, but I just couldn't get any of the Linux tools to sound as good or work as fast as Google TTS services. Also, the TTS paid services I found were just too expensive to justify (20hr book for ~$70).
I'm more than happy to additionally purchase the audiobook when it is published. I just don't want to wait.
https://beta.elevenlabs.io/speech-synthesis
is vastly better, especially for fiction.
Also worth trying is: https://speechify.com/
I don’t want to wait for the publisher to decide they want to do an audiobook.
For what its worth, most of the cost of audiobooks doesn't come from paying talent. For intermediate level actors, the going rate is around $50-$100 per finished hour (PFH) and experienced actors it can be around $250-$300. This page does a decent job of laying out pay structures for audiobooks: https://speechify.com/blog/whats-the-meaning-of-per-finished...
An 8 hour audio book might cost the author/producer about $1800-$2k.
Just talking about Audible exclusively, they take about %50 of sales. But it's kinda wishy washy about exactly how much an author will earn in royalties. It's not as much as you might think. Good article from an author here that lays out some sales numbers: https://selfpublishingadvice.org/how-audiobook-authors-are-p...
The other way that a narrator can get paid is called royalty share. That means the author/producer doesn't pay the narrator anything up front and the voice actor then relies on a small percent of each book sale to get paid. Theoretically, if an audiobook ends up really taking off then the narrator potentially could make a lot of money. But that rarely happens. Most audiobooks that you find on Audible have very, very low sales volumes.
To sum it up, it doesn't occur to lost of audiobook fans but voice acting is a very competitive industry. It takes a lot of work to make a name for yourself, and even then the most successful actors probably aren't making much more than a highly paid software engineer. For most wannabe voice actors (including myself), its something you do more for love than necessarily to make a career out of it. Though of course, lots of people do but not the majority.
This is all why I'm personally not a fan of these voice generation models. It's going to eventually make this niche industry non-competitive for real humans except for the talent that is already established. People keep blaming the actors as being too expensive when most are barely making it without secondary jobs.
Voice acting seems to be really bad career, so eliminating that job is desired, if you can deliver same quality/better product for cheaper to customers, without requiring employees to be underpaid.
I know it sucks for people in that industry, but technical progress always eliminates jobs. Calculator used to be a job, now it’s a device.
> if you can deliver same quality/better product for cheaper to customers, without requiring employees to be underpaid.
This almost never happens. Cheaper? Yes. Same or better quality? Not a chance. Automated solutions tend to allow reducing quality way below what humans workers would want to do, or even could cheaply and reliably (i.e. doing worse job than a careless one takes actual effort/skill). Like with every other case of automation replacing humans, expect the quality to be pushed down to minimum tolerable levels, as this is the point that maximizes revenue.
There’s tons of things where quality improved immensely due to automation. Engines, drugs, batteries, just to name few.
Or, put another way:
> Removing humans from the loop often directly leads to improved quality.
Yes, but improved quality for the same price means leaving money on the table, so approximately every business immediately drops quality to the baseline and pockets the difference - and from that point on, competition will optimize the quality further down.
> Of course, those are same humans that make decisions about how to use automation, so it’s not a panacea.
It's not the humans being replaced that make that decision - it's their bosses, who rent or own the automation, that make this call.
> There’s tons of things where quality improved immensely due to automation. Engines, drugs, batteries, just to name few.
Sorta, kinda. In areas with strict regulatory standards? Yes. In areas where automation improves both cost and quality, and the competitive pressure isn't very strong? Sure. With products not yet commoditized? Often enough. When it enables market segmentation? Of course.
But then you have commodities, or automation replacing people directly on the "critical path" of value chain. That's where products and services go to shit. Bonus points if automation allows to engage customers in "self-service" - i.e. outsource work to the customers.
Case in point: automated checkout machines in stores. They reduce jobs, but in theory, they could reduce queues, increase throughput, and make shopping more pleasant - win-win deal for everyone - even the cashiers could be shifted to oversight/support jobs, ensuring increased throughput and more profit for the store, for the same number of employees.
In practice, it turns out the optimal setup for the store is deploying way too few machines, and instead of having dedicated employees for oversight/support of the machines, those responsibilities are just tacked on to the workload of the existing (reduced) work force. As a result, queues are longer, customers are frustrated, overall shopping experience is shit - but the store knows perfectly well the customers will endure it anyway[1].
The market optimizes for profits, not quality or happiness. It's not just greed - money is the lifeblood of companies, and without it they die. As a result, however, competitive pressure ensures that any value or virtue that can be sacrificed to improve profits, will be sacrificed. Those who refuse get outcompeted by those who make that sacrifice. The ratchet turns, and the sacrificed value is lost forever.
--
[0] - There are many limiting factors. If the business is pushing down quality of human work too hard, they'll eventually have to deal with employee frustration, or hit limits imposed by OSHA or labor law, or just a soft limit where producing a fixed amount of goods/services costs X in labor, and there's no point in trying to save 0.1X on quality if it requires workers to put effort, which will make them produce less per unit of time, or increase variability of output, or both.
[1] - There are many reasons for it, including customers being price sensitive to the point of irrationality, usually valuing their free time at 0, and being easy to confuse with constant churn of deals. Stores also know that frustration is a fleeting feeling, while well-crafted product selection makes a store/chain sticky. Notice how automated checkout machines tend to proliferate in grocery stores and drogeries, and are seldom seen anywhere else: that's because they work best in places where customers are susceptible to factors I described earlier - and thus will endure bad experience and still come back for more. It's not like there are alternatives - competitive pressure ensures all competitors offer equally shitty experience. The ratchet made a turn, there is no going back.
I don't know where you live, but I've never seen automated checkout machines. I only have seen self checkout machines. It requires the customer to do the cashier' job and that's all.
The only reason it's not good is that it's not automated enough (if at all -- for me the self checkout machine is literally zero automation more than a regular cashier)
> The only reason it's not good is that it's not automated enough (if at all -- for me the self checkout machine is literally zero automation more than a regular cashier)
That's the point. But you are not the buyer of that automation, the store is. That automation displaced human cashiers and lowered the quality of service for customers, while generating better margins for the store (promptly eaten by competition). From your POV, i.e. customer's POV, it's not automated enough - but it's not going to be for quite a while, because there is no incentive to do it. The store doesn't stand to benefit much from additional automation, not enough to justify investment. Whether or not customers like it is irrelevant, as long as they're still coming in anyway.
You can't find an illustrator who could "cheaply and reliably" do illustrations at Midjourney's level. You just can't. If you could you would have been the biggest contractor company in the world long time ago.
But what can you do? Every black box labeled "commercial commissioned art" is now returning similarly off images, almost but not quite there. They all dropped their prices a little, so there's that - while the few black boxes offering the quality that used to be normal now cost 2-3x of what used to be normal. Hard pass.
(Meanwhile, people operating the black boxes - i.e. companies or in-house departments churning out commercial graphics cheaply - are swimming in money made on firing all their minimum-wage artists, replacing them with Midjourney or SD, and pocketing the difference. Sure, they had to drop the prices a little to clear out remaining human-powered competitors, and they will have to drop them way further once the competition restarts in the earnest - but for a short moment, they all get to make small fortunes on selling shit output, that's 100+x cheaper to produce, at roughly the same price as mediocre one before.)
Can AI be used to generate much higher quality at the same cost as human art? Sure - you'll need to spend what you used to pay an artist, whom you just fired, on generating variants and a (cheaper, at least per unit of output) human select best ones - but yes, AI can give you much better quality for the same price. But AI can also give you same quality as before for cheaper, or somewhat worse quality for much cheaper. Which is the best option to choose?
The answer, I claim, is that there is no choice - competitive pressure will force everyone to go for shittiest quality the market can bear, sold almost at cost. This will satisfy enough demand that "standard quality" offering becomes something very expensive or outright unavailable, as economics of using minimum-wage factory artists suddenly stops working.
And of course, nothing stops you from paying what we pay now for human voice actors if there continues to be a quality differential that customers care about. (Though perhaps Baumol's Cost Disease would push the price up for today's human-generated quality.)
Extrapolating further -- if the commoditized version of audio books is AI generated voice, perhaps the new job for voice actors is human narrating/acting of AI-generated content for personalized stories ('Ractives from Stephenson's book "The Diamond Age"). Who knows, human voice actors could become more in demand, not less. To be clear I wouldn't forecast this as the most likely outcome, just pointing out that there are many possible outcomes.
That's the thing though - it's not as empowering as it seems longer-term, because the "good enough" quickly drops to "barely fit for purpose"/"if it were any worse, it would be illegal to market or sell". This has been the case with most established classes of products I can think of, including pretty much anything that's been fully commoditized.
And so
> nothing stops you from paying what we pay now for human voice actors if there continues to be a quality differential that customers care about. (Though perhaps Baumol's Cost Disease would push the price up for today's human-generated quality.)
Nothing stops me today. But even if the quality differential exists, the dropping price on the low-quality version will reduce demand on the moderate-quality version, pushing its prices up and reducing number of suppliers (here, voice actors). The end result seems to always be a bifurcation: there is not enough demand to sustain a business doing decent quality work for a reasonable price, so all companies move to providing either low quality work cheaply, or high quality work at a hefty premium. The middle disappears.
In the specific context of this thread, the middle in question is the current quality of audiobooks with voice acting. The quality level available to most consumers will be below that, and the next step up will be niche recordings at high cost.
Usually on balance this falls somewhere in between -- more value for less money for the consumer, and more profit on each marginal unit of production for the producer, which is how technology progresses across most consumer goods.
This opens up for non-signed authors to release audio books.
Unless you're talking about illegal niche, in which fair but I highly doubt stores are going to accept those. All generation helps with is content that will be free.
I wrote a quick python script to read an ebook using coqui and the end result sounds pretty good. It's come in especially handy for books I want to listen to while doing yard work and stuff around the house.
Use text to speech and chatgpt to tag the character text and timestamps.
Then use a speech to speech to change the character voices or even the whole reader.
But as a product I feel like theres some legal hurdles to figure out.
This is a huge boon for independent authors, until AIs replace us as well :-) .
Things I have learned:
* A good human narrator could do much, much better, but the quality obtained this way is not totally terrible.
* The possibility to produce a section in a matter of minutes is a huge plus. The thing with a book is that it's never totally finished. If you discover a problem after you have submitted your text to a human narrator and paid $ XXXX, there is nothing you can do.
* Currently, there is no platform that I know of distributing and selling books like this. Audible only accepts audiobooks narrated by humans. To my knowledge, platforms that accept ebooks don't handle epub with media overlays. Well, Apple Books say they do but I haven't gotten it to work. There are no alternative platforms for audiobooks that I know of, but I haven't done a ton of research there.
* The possibility to have more control over emotions expressed in the speech could be a bonus, particularly for small, overly dramatic parts of the narration. Coqui TTS new editor is a step in the right direction, but their TTS doesn't sound yet as good as Elevenlabs. Voicebox seems promising, but there is no way to use it at least for now.
* Cost is a big deal 1/3. With my scripts, I pay almost nothing when I fix a typo, since most of the audio is stored in little bits in the database, and only what changes is submitted to the API. But the human time of a narrator costs much more, as it should.
* Cost is a big deal 2/3. As a reader, I have learned that how much a book sells tells me nothing about how much I will like it. But only books that have a potential to sell can afford audiobooks. If I want to listen to a story too quirky to be mainstream, or from an independent author that I follow in Twitter, the chances I'll find it as audiobook are next to none.
* Cost is a big deal 3/3. Voice narration is not the only aspect one needs to pay for. A good story needs an army of editors, proofreaders, and designers. Generally, the more an author or a publisher needs to disburse on those, the more bland and mainstream the book must become to sell and justify the investment.
-----------------------------------
Note that this is a WIP. Book chapter with automatic narration:
An epub with media overlays. It requires an epub reader that supports that standard feature of the epub 3 specification. Currently, and that I know of, there is Thorium and BookFusion for iOS.
https://drive.google.com/file/d/1U8XUB9xhu86JuketGH5WchM0obN...
An MP3 track from the epub above:
https://drive.google.com/file/d/1-u89ee52VZzGZ0oTGC_az5Uqbfs...
Here's a bunch of results on YouTube and some are really good
https://www.youtube.com/results?search_query=mariah+carey+ai...
Hopefully it’ll do a LLaMA.
Eleven Labs is the first voice synthesis that is good enough that I'd listen to an audiobook generated from it, but pricing is such that it would cost $100 to synthesize a 10 hour audiobook. A little too expensive. If they could get it down to $10 I'd cancel my Audible subscription and just synthesize audio from ebook text.
So if I can get a locally running voicebox model and just leave it running on my laptop over night transcribing an audiobook, that's even better.
Although it wasn't clear to me how voicebox compares.
1 - Is a real pain to get 'working right' - it's not even remotely batteries included
and, more importantly:
2 - Is incredibly slow. I've been turning Heart Of Darkness into an audiobook as a unit test and it takes ~30m per paragraph, on average. Add to that the occasional hiccup where a block gets transcribed badly (Tortoise occasionally 'drops out' of it's selected voice) and Tortoise only really works if you have a ton of compute and you still don't mind waiting forever.
This is basically my dream for local AI... locals models trained on my own data/code/styles. Even if they're slow, as long as they work (V/RAM) and are of high enough quality then I'm happy to wait!
They are not releasing the model (yet?) but demo samples are available across many tasks
> narrator for the presentation is indian woman with lisp
every time
https://github.com/jmiskovic/voicebox
If I ever start selling scrapbooks for collecting human faces I'll be returning the favor.
hehe, fun world we live in.
...not been open sourced and cannot be found.
Sorry AI bros. Better read the paper this time.
Yeah... I mean there definitely are ways to misuse this (especially the style transfer!) I don't think you're going to do anything except delay the inevitable Facebook.
Everyone one of their products is just garbage to me and becoming less relevant by the day. When do they actually starting building something useful again ?
Honestly Apple seems to be using “AI” much more successfully and actually seamlessly integrating it into their existing products to improve them.
My theory is Mark is hoping the meta verse will pop out if Yan’s bottom at some stage. Maybe he is right? I just can’t for the life of my understand why the current products are just so so neglected?
My Apple products continue to improve my day to day immensely. The meta products are just rubbish on the whole. How has Instagram improved in the last 5 years ?