NaturalSpeech: End-to-end text to speech synthesis with human-level quality
speechresearch.github.io
speechresearch.github.io
And even when evaluating these samples as a group, I may be imagining the distinctions I am drawing from a relatively small selection that might be cherry-picked. Nevertheless:
The generated samples are more consistent as a group, and more even in quality, with few instances of emphasis that seem (however slightly) out of place.
The recorded human samples vary more between samples (by which I mean the sample as a whole may be emphasized with a bit of extra stress or a small raising or lowering of tone compared to the other samples), and within the sample there is a bit more emphasis on a word or two or slight variance in the length of pauses, mostly appropriate in context (as in, it is similar to what I, a non-professional[0], would have emphasized if I were being asked to record these).
In general for a non-dramatic voiceover you want to maintain consistency between passages (especially if they may be heard out of order) without completely flattening the in-passage variation, but tastes vary.
Conclusion: For many types of voice work, these generated samples are comparable in quality or slightly superior to recordings of an average professional. For semi-dramatic contexts (eg. audiobooks) the generated samples are firmly in the "more than good enough" zone, more or less comparable to a typical narrator who doesn't "act" as part of their reading.
[0] Decades ago in Los Angeles I tried my hand at voiceover and voice acting work, but gave up when it quickly became clear that being even slightly prone to stuffy noses, tonsillitis and sore throats was going to pose a major obstacle to being considered reliable unless I was willing to regularly use decongestants, expectorants, and the like.
The synthesis still doesn't know where to place emphasis. You may be unable to distinguish between POOR human voice work and TTS, but not GOOD human reading.
Not a big improvement over some existing TTS. Try e.g. Michelle at: https://cloudpolly.berkine.space/
Whether that makes a big difference in practice, I don't know.
True. And yet, these samples aren't in a monotone! That's an enormous improvement.
> You may be unable to distinguish between POOR human voice work and TTS, but not GOOD human reading.
I think these are indistinguishable from AVERAGE human voice work. Keep in mind that the POOR voice work you may have in mind is probably still being done by someone who is, at least nominally, a paid professional.
*> Try e.g. Michelle at: https://cloudpolly.berkine.space/
I'm not able to select a voice other than Oscar (which is definitely worse than this).
https://www.eff.org/deeplinks/2005/09/authors-guild-sues-goo...
https://www.cnn.com/2011/09/13/tech/web/authors-guild-book-d...
https://popula.com/2022/01/22/what-kind-of-writer-accuses-li...
The difference between you reading and the software is one is 'mechanical', whether you want to constrain someone doing that in the rights you grant is debatable.
Copyright enables you (gives the creator the 'freedom' to choose) to make such choices, its what people choose to limit is the problem, as they tend to be very stingy.
The "outcome" (text is read aloud) is the same if you read it aloud to yourself. Really the difference is that when a publisher releases an audiobook they hire someone (sometimes multiple someones) often an actor or the original author to sit in a recording studio and recite. They pay for things like studio time, sound engineering, editing, the narrators time, sometimes music or foley, etc. It's very much a different product.
If you have a book and you recite it, or if you pay someone to come into your home and read it to you, or you get a bunch of software and have your computer read it for you, that's your right and at your own expense in terms of time, money, and effort. Some kind of text to speech software is expected on pretty much every device. Including such features in devices or using those features (especially the accessibility features) of your own devices isn't copyright infringement, shouldn't open you up to demands for payments from publishers, and is in no way comparable to a professionally produced audiobook. Maybe one day the tech will advance to where a program can gather the context needed to speak with and convey the correct emotion for each line and will be capable of delivering a solid performance, but right now we're lucky if more than 2/3 of the words are even pronounced correctly and the inflection isn't bizarre enough to make you question what was being said or distract you from the material.
I still expect it'll mean a lot of very amusing outbound voice mail messages.
Yeah ... so you might be getting into performance rights then ;P
As a follow up it would be cool, TTS did use different voices for Narrator and characters in a book ... if someone patents that your welcome!
Honestly, if AI ever gets good enough at crafting films from literary source material that the AI movies has any chance of competing with a hollywood production the entertainment industry is screwed. I'm positive that by then whatever crazy stuff that AI is putting out will be everywhere and playing with it would be way more fun than a movie theater ticket.
Different voices would be cool, but who is speaking which line can sometimes be ambiguous even for human readers. I'd be happy with just one voice that didn't sound like a robot or like a human voice spliced together from multiple sources.
We recently listened to How To Train Your Dragon instead of reading it to our kids ourselves specifically because it was narrated by David Tennant.
That said, having only skimmed the paper I didn’t notice a discussion of the compute requirements for usage (just training), but it did say it was a 28.7 million parameter model, so I recon this could be used in real-time on a phone.
[0] judging by the videos of Dr. Sbaitso on YouTube, it was only one step up from the intro to Impossible Mission on the Commodore 64.
Microsoft and particularly Apple used to make cutting edge (at the time) TTS available as a marketing gimmick, but then did not develop this further because TTS was only relevant for screen readers and people with impaired vision weren't really their primary customers. Then TTS made advances and companies started selling high quality voices at high price points. Now companies want to make as much money with it as they can. Improving "free" TTS for ordinary customers is not really a top priority of Microsoft and Apple. Moreover, since network speeds have increased, any improved end-consumer TTS will send the text to a server and the audio back, so that a company can collect and analyze all the texts and make money with this spy data. That's how Google's free TTS server works.
That's exactly what I meant when I wrote "...strive to keep the status quo of being way beyond the capabilities of the people".
Monopolizing the means of maximizing the profits to maximize the power - now that's some deep conspiracy territory, indeed.
Honestly, are those really different? How is it not in their interest for money?
But really, the situation is pretty good, with a lot of code and dataset available as opensource. Notably, if you're not constrained to smartphones and the like, you can run on your computer quite a number of modern models, see for instance https://github.com/coqui-ai/TTS/ (which itself contains many different models).
The work that needs to be done is """just""" to turn those models into something suitable for smartphones (which will most likely include re-training), and to plug them back into Android's TTS API.
* Larynx: https://github.com/rhasspy/larynx/
* OpenTTS: https://github.com/synesthesiam/opentts
* Likely Mimic3 in the near future: https://mycroft.ai/blog/mimic-3-preview/
Larynx in particular has a focus on "faster than real-time" while OpenTTS is an attempt to package & provide common REST API to all Free/Open Source Text To Speech systems so the FLOSS ecosystem can build on previous work supported by short-lived business interests, rather than start from scratch every time.
AIUI the developer of the first two projects now works for Mycroft AI & is involved in the development of Mimic3 which seems very promising given how much of an impact on quality his solo work has had in just the past couple of years or so.
Here's another project that builds on top of VOSK to provide a tighter integration with Linux: https://github.com/ideasman42/nerd-dictation
One of the vast litany of non-AI related Siri flaws that make you wonder how a $2t company can achieve such tragic levels of incompetence.
Huh.
For example, compare "rebuke and abash": in the NaturalSpeech, one goes down like she's sure and the other goes up like she's questioning, where in the recording, they are both more balanced and emphasized as equally important words in the sentence. And the pause after insolent in "insolent and daring" sounds uneven compared to the recording, which emphasizes the pair of words more equally and tightly.
Jiminy Glick interviews (and does an impression of) Jerry Seinfeld:
It would get really messy of you had to put tags around individual letters, and letters don't even map directly to syllables, so you'd need to mark up a phonetic transcription. At that point you might as well use a binary file format, not xml.
But for that kind of stuff (singing), there are great tools like Vocaloid (which is a HUGE thing in Japan):
https://en.wikipedia.org/wiki/Vocaloid
VOCALOID5 - Walkthrough
https://www.youtube.com/watch?v=UAtVGHl1AFM
(Check out the "Cool / Cute" slider at 8:22!)
Here's a much simpler and cruder tool I made years ago (when xml was all the rage) for editing and "rotoscoping" speech pitch envelopes, called the "Phoneloper" -- not quite as polished and refined as Vocaloid, but it was sure fun to make and play with:
Phoneloper Demo
The Phoneloper is a toy/tool for creating and editing expressive speech "Phonelopes" that Don Hopkins developed for Will Wright's Stupid Fun Club in 2003, using Python + Tkiter + CMU's "Flite" open source speech synthesizer. It modified Flite so it could export and import the diphone/pitch/timing as xml "Phonelopes", so you could synthesize a sentence to get an initial stream of diphones and a pitch envelope. Then you could edit them by selecting and dragging them around, add and delete control points from the pitch and amplitude tracks to inflect the speech, stretch the diphones to change their duration, etc. It was not "fully automatic", but you could load an audio file and draw its spectrogram in the background of the pitch track, so you stretch the diphones and "rotoscope" the pitch track to match it.
There's this: http://zeehio.github.io/festival/doc/Singing-Synthesis.html It won't sound as good as Vocaloid, but it uses XML with singing synthesis.
This is the example code to make it sing Happy Birthday:
>DECtalk Software Sample Singing Program
[:phoneme on]
[hxae<300,10>piy<300,10> brr<600,12>th<100>dey<600,10> tuw<600,15> yu<1200,14>_<120>] [hxae<300,10>piy<300,10> brr<600,12>th<100>dey<600,10> tuw<600,17> yu<1200,15>_<120>] [hxae<300,10>piy<300,10> brr<600,22>th<100>dey<600,19>dih<600,15>rdeh<600,14&g;ktao<600,12>k_<120>_<120>] [hxae<300,20>piy<300,20> brr<600,19>th<100>dey<600,15> tuw<600,17> yu<1200,15>]
---
[1] https://vt100.net/manx/details/1,1230 (92mb pdf so I won't link directly to it, but info on singing feature appears on p. 116 for the curious)
Eedie & Eddie (And The Reggaebots) - Some Velvet Morning (Peter Langston)
https://www.youtube.com/watch?v=1l0Ko1GUiSo
Peter S. Langston - "Some Velvet Morning" (By Lee Hazelwood) - Performed By Eedie & Eddie And The Reggaebots
http://www.wfmu.org/365/2003/169.shtml
Eedie & Eddie On The Wire
http://www.langston.com/SVM.html
Peter Langston's Home Page:
His 1986 Usenix "2332" paper:
http://www.langston.com/Papers/2332.pdf
How to use Eddie and Eedie to make free third party long distance phone calls (it's OK, Bellcore had as much free long distance phone service as they wanted to give away for free):
https://news.ycombinator.com/item?id=22308781
>My mom refused to get touch-tone service, in the hopes of preventing me from becoming a phone phreak. But I had my touch-tone-enabled friends touch-tone me MCI codes and phone numbers I wanted to call over the phone, and recorded them on a cassette tape recorder, which I could then play back, with the cassette player's mic and speaker cable wired directly into the phone speaker and mic.
>Finally there was one long distance service that used speech recognition to dial numbers! It would repeat groups of 3 or 4 digits you spoke, and ask you to verify they were correct with yes or no. If you said no, it would speak each digit back and ask you to verify it: Was the first number 7? ...
>The most satisfying way I ever made a free phone call was at the expense of Bell Communications Research (who were up to their ears swimming in as much free phone service as they possibly could give away, so it didn't hurt anyone -- and it was actually with their explicitly spoken consent), and was due to in-band signaling of billing authorization:
When you called (201) 644-2332, it would answer, say "Hello," pause long enough to let the operator ask "Will you accept a collect call from Richard Nixon?", then it would say "Yes operator, I will accept the charges." And that worked just fine for third party calls too!
>Peter Langston (working at Bellcore) created and wrote a classic 1985 Usenix paper about "Eedie & Eddie", whose phone number still rings a bell (in my head at least, since I called it so often): [...]
>(201) 644-2332 or Eedie & Eddie on the Wire: An Experiment in Music Generation. Peter S Langston. Bell communications Research, Morristown, New Jersey.
>ABSTRACT: At Bell Communications Research a set of programs running on loosely coupled Unix systems equipped with unusual peripherals forms a setting in which ideas about music may be "aired". This paper describes the hardware and software components of a short automated music concert that is available through the public switched telephone network. Three methods of algorithmic music generation are described.
Of course, whether a particular speech synthesis system supports such features is another thing.
It also has a `pitch_contour` attribute: https://www.w3.org/TR/speech-synthesis11/#pitch_contour
Coincidentally enough I pretty much only know any of this because a few years ago I created a GUI for a client which enabled an assistive technology researcher to "draw" in the pitch contour required for a word/phrase (not unlike the project demonstrated in your video :) ) from which the SSML was then generated.
I tend to view most of these things through the perspective of what would help mod-maker's for video games: VA is the one thing you basically can't do yourself, but you also tend to have a decent data set to pull from (and I suspect various open source voice sample sets would become pretty popular).
There's definitely research/proprietary software that can enable a person speaking in desired manner to have their voice control the expression of the generated speech.
Here's a related issue on a Open Source text to speech project which I only learned of today: https://github.com/neonbjb/tortoise-tts/issues/34#issue-1229...
> I tend to view most of these things through the perspective of what would help mod-maker's for video games
Yeah, I think there's some really cool potential for indie creatives to have access to (even lower quality) voice simulation--for use in everything from the initial writing process (I find it quite interesting how engaging it is to hear one's words if that's going to be the final form--and even synthesis artifacts can prompt an emotion or thought to develop); to placeholder audio; and, even final audio in some cases.
> (and I suspect various open source voice sample sets would become pretty popular).
That's definitely a powerful enabler for Free/Open Source speech systems. There's a list of current data sets for speech at the "Open Speech and Language Resources" site: https://openslr.org/resources.php
Encouraging people to provide their voice for Public Domain/Open Source use does come with some ethical aspects that I think people need to be made aware of so they can make informed decisions about it.
Given your interest in this topic you might be interested in this (rough) tool I finally released last week: https://rancidbacon.itch.io/dialogue-tool-for-larynx-text-to...
And sometimes even the director doesn't notice. Even when they also wrote the screenplay! My favourite example is from The Matrix, @1:01:40, where Cypher says:
> The image translators work FOR the construct program, but there's way too much information to decode the Matrix
Stress on the "for", which makes no sense; he sounds like he's revealing an employee/employer relationship. The stress should be on the start of "CONstruct", since he's saying the tech works for that but not for the other. Subtle, but it changes the whole sense of the line.
Goddamn you Cypher, indeed.
Whether it's supported by a particular speech synthesis system and which set of features are supported varies.
The standard does include a section on "3.2 Prosody and Style" which covers this aspect to some degree: https://www.w3.org/TR/speech-synthesis11/#S3.2
This isn't exactly it but it's very close - https://www.deepmind.com/blog/wavenet-a-generative-model-for... CTRL+F 'babbling'
Katy Perry - Last Friday Night in Simlish:
https://www.youtube.com/watch?v=sxyW6AJ-yIk
How the Language From the Sims Was Created:
https://www.youtube.com/watch?v=FGsbeTV76YI
Simlish Voice Video (Gerri Lawlor and Stephen Kearin, inventors of Simlish):
https://www.youtube.com/watch?v=Y_E6026i9tA
Steve and Gerri improvising together in English while playing The Sims:
https://donhopkins.com/home/catalog/sounds/Steve_And_Gerri.w...
Gerri Lawlor:
https://en.wikipedia.org/wiki/Gerri_Lawlor
At the University of Maryland VAX Lab in the 80's, we had a DECTalk attached to the VAX over a serial line that we'd play around with, but I think the protocol must have used two byte tokens, because some times it would get one byte out of sync and start going "BLLEEGH YAAUGH RAWGH BRAGHK SPROP BLOP BLOP GUKGUK BWAUGHK GYAAUGHT BLOBBLE SPLOP BLAP BLAP BEAUGH GUWK SPLAPPLE PLAP SPLORPLE BLAPPLE"! (*)
Just like it was channeling the Don Martin Sound Effects from random Mad Magazines.
https://www.madcoversite.com/dmd-alphabetical.html
(*) Bulemia Meeting Attendees Vomiting, MAD #266 1987, Page 45, On Thursday Evening on West 12th Street.
https://www.thorsten-voice.de/
https://github.com/thorstenMueller/Thorsten-Voice
where someone contributed a huge set of his voice samples and a tutorial / script collection to build a pretty decent TTS model LOCALLY.
Quality-wise it is not as good as the samples in the article, but its free and pretty easy to follow for a tech enthusiast.
> Authors
> ...
> Microsoft Research Asia & Microsoft Azure Speech
https://azure.microsoft.com/en-us/blog/announcing-new-voices...
My opinion is Microsoft's Azure TTS is the best and IBM's Watson TTS was a pretty close second. I remember finding Google's disappointing as well.
https://azure.microsoft.com/en-us/services/cognitive-service...
I’d love to see something that could actually do choir synthesis using a method like the one in the article.
But here's I guess a more pure-synth version (depending on what that means, like is any real voice data OK at any point in the algo?), not too bad.
https://youtu.be/9vEA4iSjajg?t=517
Edit: Ah, here's a type & speak version...who knows what cheats there are, could be somebody under his desk for all I know:
https://www.youtube.com/watch?v=NNyQ7FWV2E8
Less drama, though different software:
Command line is fine but it would be much better if it could trivially take clipboard content for input. The last time I looked I found stuff that wasn't that great and was pretty inconvenient.
* Larynx: https://github.com/rhasspy/larynx/
* OpenTTS: https://github.com/synesthesiam/opentts
* Likely Mimic3 in the near future: https://mycroft.ai/blog/mimic-3-preview/
Both the NaturalSpeech and the human said pretty much every word in that sentence completely incorrectly for the context of the words. It is the difference between "the car Seat" and "the car seat". "It's pronounced Ore-garh-no" to paraphrase the insufferable Hermione Granger.
There is a rather obvious problem with the stress on "warehouses", and a more subtle problem with "warrants on them", where it's difficult to get the stress pattern just right.
Sounds really good though.
I created these samples in a relatively short time using the Free/Open Source (which I think is an important factor for indies) text-to-speech project Larynx & an narrative editor I finally released the other weekend:
* https://github.com/rhasspy/larynx/
* https://rancidbacon.itch.io/dialogue-tool-for-larynx-text-to...
Now, I would really like to link you directly to audio of the next two but considering it's currently in beta behind an (automated response) email address, I think that may not be appropriate, so, instead...
* Visit & get access to the beta here: https://mycroft.ai/blog/mimic-3-preview/
* Copy & paste this SSML into the form: https://pastebin.com/Bwd7LCbj
It's definitely a noticeable step up again in quality.
There's an alternate pair of voices if you move the "_" from one "name" attribute to the other in each "voice" element.
I intentionally didn't edit the text to remove some of the artifacts both to give a realistic impression of the current state & because sometimes they add interesting texture. :)
Note the beta voices are "low" quality.
Sounds like openly reproducing this result is within independent researchers’ reach.
You can click play on any/all of the samples simultaneously, resulting in a neat sonic effect vaguely reminiscent of Steve Reich's famous "Come out." [1]
[1] https://www.youtube.com/watch?v=g0WVh1D0N50 (skip to like 7 minutes in to get the idea)
Why don't they use this tech to recreate some dead actor's speech, for example?
edit; is the Nuance acquisition compounding yet?
What's the end game here? because I cannot use it, I cannot buy it and this seems more than just a scientific paper.
So what's the objective here?
I'd love to be able to use an assistant device in Croatian in my lifetime.
I recently tried out the open source android TTS engines and they don't seem all that great even though there's been years of development on them.
Can anyone that knows comment on what the complexity here is?
* https://rhasspy.github.io/larynx/
* https://mycroft.ai/blog/mimic-3-preview/ (forthcoming, currently in beta, actively looking for feedback on non-English language quality)
From a quick look here, there didn't seem to be any Croatian open source data options either:
* https://openslr.org/resources.php
I have a vague recollection that there was at least one (I think) Eastern European country that managed to get government funding in order to support creation of local language assistive device text to speech.
So, it doesn't seem like an impossible task but certainly a non-zero amount of work to collect & process appropriate audio data.
Hope you get your dream at some point in future. (As everyone deserves assistive devices in their own language.)
Most of Ex-Yugoslavian languages (The southwest Slavic group) are basically the same. Much along the lines of US vs UK English. Slovenian and Macedonian are also quite understandable.
Hopefully advances in AI will allow for a more general approach which will accelerate the development of these technologies.
This is on contrast to approaches that involve pre- and postprocessing (e.g. sending pronunciation tokens to the model, or models returning FFT packets or TTS parameters instead of raw waveforms).
It's a common and well understood technical term in this context.