ESpeak-ng: speech synthesizer with more than one hundred languages and accents
github.com
github.com
Join that with algorithms which can translate English into phonetic tokens with relatively high accuracy, and you have speech synthesis. Make the dictionary big enough, add enough finesse, and a few hundred rules about transitioning from phoneme to phoneme, and it's produces relatively understandable speech.
Part of me feels that we are losing something, moving away from these classic approaches to AI. It used to be that, to teach a machine how to speak, or translate, the designer of the system had to understand how language worked. Sometimes these models percolated back into broader thinking about language. Formant synthesis ended up being an inspiration to some ideas for how the brain recognizes phonemes. (Or maybe that worked in both directions.) It was thought, further advances would come from better theories about language, better abstractions. Deep learning has produced far better systems than the classic approach, but they also offer little in terms of understanding or simplifying.
I don't know of any way to go smaller than that with software. I tried, but it seems like a fundamental limit for English.
If you include "robotic" speech, then there's https://en.wikipedia.org/wiki/Software_Automatic_Mouth in a few tens of KB, and the demoscene has done similar in around 1/10th that. All formant synths, of course, not the sample-based ones that you're referring to.
That is, instead of allowing a generic machine learning model to output unconstrained audio, train it on the basis of letting it produce low bitrate input/control values for a formant synth instead, and see just how small you can push the model.
For phonetically simple languages, such a system can easily fit on a microcontroller with kilobytes of RAM and a slow CPU. English might require a little bit more on the text-to-phoneme stage, but you can definitely go far below 1MB.
For the CMU flite voices they represent the data as LPC (linear predictive coding) data with residual remainder (residual excited LPC). The HTS models use simple neural networks to predict the waveforms -- IIRC, these are similar to RNNs.
The MBROLA models use OLA (overlapped add) to overlap small waveform samples. They also use diphone samples taken from midpoint to midpoint in order to create better phoneme transitions.
That's a descendant of Festival Singer, which was well respected in its day.
What's a current practical text-to-speech system that's open source, local, and not huge?
The size of the associated voice files varies but there are options that are under 100MB: https://huggingface.co/rhasspy/piper-voices/tree/main/en
It does feel like we're rapidly losing this relationship in general. I think it's going to be a good thing overall for productivity and the advancement of mankind, but it definitely takes a lot of the humanity out of our collective accomplishments. I feel warm and fuzzy when a person goes on a quest to deeply understand a subject and then shares the fruits of their efforts with everyone else, but I don't feel like that when someone points at a subject and says "hey computer, become good at that" with similar end results.
I think AI will only cause us to become stuck in another local maximum, since not understanding how something works can only lead to imitation at best, and not inspiration.
E.g. we know human voices can be produced well with formant synthesis because we know how the human vocal tract is shaped. So you can "give" a model a formant synth, and try to train smaller models outputting to it.
I think there's going to be a whole lot of research possibilities in placing constraints and training smaller models, and even training ensembles of models constrained in how they're interacting and their relative sizes to try to "force" extraction of functionality.
E.g. we have reasonable estimates at the lowest bitrate raw audio that produces passable voice. Now consider training two models A, B, where A => B => audio, and the "channel" between A and B is constrained to a small fraction of the bitrate that'd let A do all the work, and where the size of B is set at a level you've first struggled to get passable TTS output from.
Try to squeeze the bitrate and/or the size of B down and see if you can get something to emerge where analysing what happens in B is doable.
One of the things I've long wanted to do but not found time for, is to take a few different variants of formant synths, and try to train a simple TTS model to control one instead of producing "raw" output. It's amazing what TTS models can do with "raw" output, but we know our brains aren't producing raw, unconstrained digital audio, and so I think there's a lot of potential to understanding more and simplifying if you train models constrained to produce outputs we know ought to be sufficient, and push their size as far down as we can.
The videos on the other link does show that the guy who has done that API has done stuff to the UI to allow interfacing to it in non-graphical ways, but sadly I don't see any online demos of those alternative user interfaces anywhere, which is a great shame. Sadly it doesn't look like the videos are very accessible either, as they're mostly demos with no commentary about what is going on on the screen, without which they're mostly just random sounds.
Thanks for the link to the programmable version--I don't think I'd been aware of that previously...
[0] And, from personal experience, also rather difficult to "safely" search for if you don't quite remember its name exactly. :D
I took a few stabs at understanding Klatt, but I feel like I had far too little DSP, math and linguistic intuitions back then to fully comprehend it, perhaps I should take another one now.
Something similar from the 800s is the Euphonia talking machine ( https://en.m.wikipedia.org/wiki/Euphonia_(device) ).
Clicked that thinking someone had made a talking machine in the Middle Ages :)
Absolutely. Seems a large amount of software developers have moved on from trying to understand how things work to solve the problem, and they are now instead just essentially throwing shit at a magical wall until something sticks long enough.
I depend on TTS to overcome dyslexia, but I also struggle with auditory processing disorder that causes me to misunderstand words. As a result , classical TTS does not help me read faster or more accurately than struggling through my dyslexia. It causes me to rapidly fatigue, zone out, and rewind often, in a way that is more severe than when I sight read.
On the other hand, modern neural TTS is a huge enabler. My error rate, rewind rate, and fatigue are much better thanks to the natural tone, articulation, and prosody. I'm able to read for hours this way, and my productivity is higher than sight reading alone. This unlocks long and complex readings that I would never complete by sight reading alone, like papers in history, philosophy, and law. Previously I was limited to reading math, computer science, and engineering work, where I heavily depended on diagrams and math formulas to help me gloss over dense text readings.
The old tech had no impact on my life, given my combination of reading and listening difficulty, since it was not comparatively better than sight reading. But my life changed about 6 years ago with neural TTS. The improvement has been massive, and has helped me work with many non-technical readings that I would previously give up on.
The main issues I see now is not that neural models are hard to understand. For better or worse, we're able to improve the models just by throwing capital and ML PhDs at the problem. The problem I see is that the resulting technology is proprietary and not freely available for the people whose lives it would change.
We should work towards a future where people can depend on useful and free TTS that improves their quality of life. I don't think simple synthetic models will be enough. We must work to seize control of models that can provide the same quality of life improvements that new proprietary models can provide. And we must make these models free for everyone to use!
There are also plenty of places where the current modern "neural" models are too compute intensive / costly to run, and so picking just the current big models isn't an option for all uses.
I switched to it in early childhood, at a time where human-sounding synthesizers were notoriously slow and noticeably unresponsive, and just haven't found anything better ever since. I've used Vocalizer for a while, which is what iOS and Mac OS ship with, but then third-party synthesizer support was added and I switched right back.
I tried a bunch of speech synthesis, with speed and intelligibility in mind.
ESpeakng-ng barely intelligible past ~500 words per minute, and just generally unpleasant to listen to. Maybe my brain just can't acclimatize to it.
Microsoft Zira Mobile (unlock on win11 desktop via regex) sounds much more natural and intelligible at max windows SAPI speech rate, which I estimate is around ~600 and equivalent to most conversation/casual spoken word at 2x speed. I wish windows could increase playback even further, my brain can process 900-1200 words per minute or 3x-4x normal playback speed.
On Android, Google's "United States - 1" sounds a little awkward but also intelligible at 3x-4x speed.
Like usual the limit isn't how fast human io is but how fast human processing works.
>text is heavy going enough I have to drop it down to 100 wpm
What is heavy text for you? Like very dense technical text?
>put it on loop
I find this very helpful as well, but for content I consume, not very technical, I listen at ~600wpm and loop it multipe times. It's like listening a song to death. Engrain it on a vocal / story telling level.
E: semi related comment to a deleted comment about processing speed that I can no longer reply to. Posting here because related.
Some speech synthesis are much more intelligible at higher speeds, and aids processing at higher wpms. What I've been trying to find is the most intelligible speech synthesis voice for upper limit of concentrated/burst listening which for me is around 1200wpm / 4x speed, i.e. many have wierd audio artefacts past 3x. There's synthesis engines whose high speed intelligbility improves if text is processed with SSML markup to add longer pauses after punctuation. Just little tweaks that makes processing easier. Doesn't apply to all content, all contexts, but I think some consumption are suitable for that, and it's something that can be trained like many mental tasks, and dedicated speech synthesis like fancy sport equipments improve top end performance.
IMO also something neural model can be tuned for. There are some podcasters/audiobook narrators who are "easy" listening at 3x speed vs others because they just have better enunciation/cadence at same word density. Most voices out there from traditional SAPI models to neural are... very mid fast "narrators". Think need to bundle speech sythensis with content awareness - AI to filter content then synthesis speech that emphasis/slow on significant information, breeze past filler - just present information more efficiently for consumption.
I've been spoiled by modern AI generated voices that sound indistinguishable from humans to me.
If this is just how it is, the "more than one hundred languages" claim is a bit suspect.
Here's some potentially related links if you'd like to dig deeper:
* "questions about mandarin data packet #1044": https://github.com/espeak-ng/espeak-ng/issues/1044
* "ESpeak NJ-1.51’s Mandarin pronunciation is corrupted #12952": https://github.com/nvaccess/nvda/issues/12952
* "The pronunciation of Mandarin Chinese using ESpeak NJ in NVDA is not normal #1028": https://github.com/espeak-ng/espeak-ng/issues/1028
* "When espeak-ng translates Chinese (cmn), IPA tone symbols are not output correctly #305": https://github.com/rhasspy/piper/issues/305
* "Please default ESpeak NG's voice role to 'Chinese (Mandarin, latin as Pinyin)' for Chinese to fix #12952 #13572": https://github.com/nvaccess/nvda/issues/13572
* "Cmn voice not correctly translated #1370": https://github.com/espeak-ng/espeak-ng/issues/1370
However, I wasn't satisfied with the speech quality so now I'm using RHVoice. RHVoice seems to produce more natural/human-sounding output yo me.
RHVoice sounds slightly more natural in some cases, but one advantage of espeak-ng is that the text parsing logic is cleaner, by default.
For example, RHVoice likes to spell a lot regular text formatting. One example would be spelling " -- " as dash-dash instead of pausing between sentences. So while text sounds a little more natural, it's actually harder to understand in context unless the text is clean to begin with.
I don't know if speech-dispatcher does this for you, but I'm using a shell script and some regex rules to make the text cleaner for TTS which I don't need when using espeak-ng.
Another tradeoff: espeak-ng with the mbrola doesn't offer all the inflexion customization options you have with the "robotic-sounding" voices. When accelerating speech, these options make a qualitative difference in my experience.
I can see why each of these can have its place.
The default voice is optimized for space and speed instead of quality of the generated audio.
It’s simply the wrong tuning tradeoff.
* accessibility
* non-accessibility (e.g. voice interfaces; narration; voice over)
The qualities of the generated speech which are favoured may differ significantly between the two domains, e.g. AIUI non-accessibility focused TTS often prioritises "realism" & "naturalness" while more accessibility focussed TTS often prioritizes clarity at high words-per-minute speech rates (which often sounds distinctly non-"realistic").
And, AIUI espeak-ng has historically been more focused on the accessibility domain.
[1] https://github.com/espeak-ng/espeak-ng/blob/master/docs/lang...
> eSpeak NG uses a "formant synthesis" method. This allows many languages to be provided in a small size. The speech is clear, and can be used at high speeds, but is not as natural or smooth as larger synthesizers which are based on human speech recordings. It also supports Klatt formant synthesis, and the ability to use MBROLA as backend speech synthesizer.
I've been using eSpeak for many years now. It's superb for resource constrained systems.
I always wondered whether it would be possible to have a semi-context aware, but not neural network, approach.
I quite like the sound of Mimic 3, but it seems to be mostly abandoned: https://github.com/MycroftAI/mimic3
IIUC ongoing development of Piper TTS is now financially supported by the recently announced Open Home Foundation (which is great news as IMO synesthesiam has almost single-handed revolutionized the quality level--in terms of naturalness/realism--of FLOSS TTS over the past few years and it would be a real loss if financial considerations stalled continued development): https://www.openhomefoundation.org/projects/ (Ok, on re-reading OHF is more generally funding development of Rhasspy of which Piper TTS is one component.)
"At the roundabout, the second exit take."
"At your destination, arrived have you."
"A star shall shine on the hour of our taking the second exit."
"You have reached your Destination, fair as the Sea and the Sun and the Snow upon the Mountain!"
Or you want a model which translates normal English into Yoda-English (on text level) and then attach a speech synthesizer on that?
Or I guess an end-to-end speech synthesizer, a big neural network which operates on the whole sentence at once, could also internally learn to do that grammar transformation.
Most times I try to use modern "natural-sounding" voices they take a while to initialize, and when you speed them at a certain point the words mix together into meaningless noise, while at the same rate eloquence and espeak would handle just great, well, for me at least.
I was thinking about this a few days back while I was trying out piper-tts [0] how supposedly "more advanced" synthesizers powered by AI use up more ram and cpu and disk space to deliver a voice which doesn't sound much better than something like RH voice and gets things like inflection wrong. And that's the english voice, the voice for my language (serbian) makes espeak sound human and according to piper-tts it's "medium".
Funny story about synthesizers taking a while to initialize, there's a local IT company here that specializes in speech synthesis and their voices take so long to load they had to say "<company> Mary is initializing..." whenever you start your screen reader or such. Was annoying but in a fun way. Their newer Serbian voices also have this "feature" where they try to pronounce some english words it comes upon properly. It also has another "feature" where it tries to pronounce words right that were spelled without accent marks or such, and like with most of these kinds of "features" they combine badly and hilariously. For example if you asked them to pronounce "topic" it would pronounce it as "topich, which was fun while browsing forums or such.
I would be very glad if there's a truly open source local hosted text to speech software which brings good human sounding speech in woman/man german/english/french/spanish/russian/arabic language...
There aren't many options for degoogled Android users. In the end I settled for the Google Speech Services and disabled network access and used the default voice. GSS has its issues and voices don't download properly, but the default voice is tolerable in this situation.
I have an Emic2 board I use (through UART so my ESP32 can send commands to it) and I use Home Assistant to send notifications to it. My family are science nerds like me, so when the voice of Stephen Hawking tells us there is someone at the door, it brings a lot of joy to us.
There are better open source TTS models. E.g. check https://github.com/neonbjb/tortoise-tts or https://github.com/NVIDIA/tacotron2. Or here for more: https://www.reddit.com/r/MachineLearning/comments/12kjof5/d_...
From the samples I listened, it sounds great to me?
They should make it so that I can do
sudo apt-get install deepspeech
sudo ln -s /usr/bin/deepspeech /usr/bin/espeak
Anything more than that is an impediment to mass adoption.Seems they need some new product management ...
For Text To Speech, I've found Piper TTS useful (for situations where "quality"=="realistic"/"natual"): https://github.com/rhasspy/piper
For Speech to Text (which AIUI DeepSpeech provided), I've had some success with Vosk: https://github.com/alphacep/vosk-api
It seems Piper currently abstracts its phonemize-related functionality with a library[0] that currently makes use of a espeak-ng fork[1].
Unfortunately it also seems license-related issues may have an impact[2] on whether Piper continues to make use of espeak-ng.
For your specific example of handling 1984 as a year, my understanding is that espeak-ng can handle situations like that via parameters/configuration but in my experience there can be unexpected interactions between different configuration/API options[6].
[0] https://github.com/rhasspy/piper-phonemize
[1] https://github.com/rhasspy/espeak-ng
[2] https://github.com/rhasspy/piper-phonemize/issues/30#issueco...
[3] Previously I've made note of some potential options here: https://gitlab.com/RancidBacon/notes_public/-/blob/main/note...
[4] For example, as I note here[5] there's currently at least four different ways to access espeak-ng's phoneme-related functionality--and it seems that they all differ in their output, sometimes consistently and other times dependent on configuration (e.g. audio output mode, spoken punctuation) and probably also input. :/
[5] https://gitlab.com/RancidBacon/floss-various-contribs/-/blob...
[6] For example, see my test cases for some other numeric-related configuration options here: https://gitlab.com/RancidBacon/floss-various-contribs/-/blob...
I use chrome's extension 'read aloud', which is as natural as you can get.
* the project seems to provide real value to a huge number of people who rely on it for reasons of accessibility (even more so for non-English languages); and,
* the project is a valuable trove of knowledge about multiple languages--collected & refined over multiple decades by both linguistic specialists and everyday speakers/readers; but...
* the project's code base is very much of "a different era" reflecting its mid-90s origins (on RISC OS, no less :) ) and a somewhat piecemeal development process over the following decades--due in part to a complex Venn diagram of skills, knowledge & familiarity required to make modifications to it.
Perhaps the prime example of the last point is that `espeak-ng` has a hand-rolled XML parser--which attempts to handle both valid & invalid SSML markup--and markup parsing is interleaved with internal language-related parsing in the code. And this is implemented in C.
[Aside: Due to this I would strongly caution against feeding "untrusted" input to espeak-ng in its current state but unfortunately that's what most people who rely on espeak-ng for accessibility purposes inevitably do while browsing the web.]
[TL;DR: More detail/repros/observations on espeak-ng issues here:
* https://gitlab.com/RancidBacon/floss-various-contribs/-/blob...
* https://gitlab.com/RancidBacon/floss-various-contribs/-/blob...
* https://gitlab.com/RancidBacon/notes_public/-/blob/main/note...
]
Contributors to the project are not unaware of the issues with the code base (which are exacerbated by the difficulty of even tracing the execution flow in order to understand how the library operates) nor that it would benefit from a significant refactoring effort.
However as is typical with such projects which greatly benefit individual humans but don't offer an opportunity to generate significant corporate financial return, a lack of developers with sufficient skill/knowledge/time to devote to a significant refactoring means a "quick workaround" for an specific individual issue is often all that can be managed.
This is often exacerbated by outdated/unclear/missing documentation.
IMO there are two contribution approaches that could help the project moving forward while requiring the least amount of specialist knowledge/experience:
* Improve visibility into the code by adding logging/tracing to make it easier to see why a particular code path gets taken.
* Integrate an existing XML parser as a "pre-processor" to ensure that only valid/"sanitized"/cleaned-up XML is passed through to the SSML parsing code--this would increase robustness/safety and facilitate future removal of XML parsing-specific workarounds from the code base (leading to less tangled control flow) and potentially future removal/replacement of the entire bespoke XML parser.
Of course, the project is not short on ideas/suggestions for how to improve the situation but, rather, direct developer contributions so... shrug
In light of this, last year when I was developing the personal project[0] which made use of a dependency that in turn used espeak-ng I wanted to try to contribute something more tangible than just "ideas" so began to write-up & create reproductions for some of the issues I encountered while using espeak-ng and at least document the current behaviour/issues I encountered.
Unfortunately while doing so I kept encountering new issues which would lead to the start of yet another round of debugging to try to understand what was happening in the new case.
Perhaps inevitably this effort eventually stalled--due to a combination of available time, a need to attempt to prioritize income generation opportunities and the downsides of living with ADHD--before I was able to share the fruits of my research. (Unfortunately I seem to be way better at discovering & root-causing bugs than I am at writing up the results...)
However I just now used the espeak-ng project being mentioned on HN as a catalyst to at least upload some of my notes/repros to a public repo (see links in TLDR section above) in that hopes that maybe they will be useful to someone who might have the time/inclination to make a more direct code contribution to the project. (Or, you know, prompt someone to offer to fund my further efforts in this area... :) )
[0] A personal project to "port" my "Dialogue Tool for Larynx Text To Speech" project[1] to use the more recent Piper TTS[2] system which makes use of espeak-ng for transforming text to phonemes.
[1] https://rancidbacon.itch.io/dialogue-tool-for-larynx-text-to... & https://gitlab.com/RancidBacon/larynx-dialogue/-/tree/featur...
[2] https://github.com/rhasspy/piper
[3] Very much no shade toward the project intended.