I don't know of any way to go smaller than that with software. I tried, but it seems like a fundamental limit for English.
I don't know of any way to go smaller than that with software. I tried, but it seems like a fundamental limit for English.
If you include "robotic" speech, then there's https://en.wikipedia.org/wiki/Software_Automatic_Mouth in a few tens of KB, and the demoscene has done similar in around 1/10th that. All formant synths, of course, not the sample-based ones that you're referring to.
That is, instead of allowing a generic machine learning model to output unconstrained audio, train it on the basis of letting it produce low bitrate input/control values for a formant synth instead, and see just how small you can push the model.
For phonetically simple languages, such a system can easily fit on a microcontroller with kilobytes of RAM and a slow CPU. English might require a little bit more on the text-to-phoneme stage, but you can definitely go far below 1MB.
For the CMU flite voices they represent the data as LPC (linear predictive coding) data with residual remainder (residual excited LPC). The HTS models use simple neural networks to predict the waveforms -- IIRC, these are similar to RNNs.
The MBROLA models use OLA (overlapped add) to overlap small waveform samples. They also use diphone samples taken from midpoint to midpoint in order to create better phoneme transitions.
That's a descendant of Festival Singer, which was well respected in its day.
What's a current practical text-to-speech system that's open source, local, and not huge?
The size of the associated voice files varies but there are options that are under 100MB: https://huggingface.co/rhasspy/piper-voices/tree/main/en