I don't know of any way to go smaller than that with software. I tried, but it seems like a fundamental limit for English.
If you include "robotic" speech, then there's https://en.wikipedia.org/wiki/Software_Automatic_Mouth in a few tens of KB, and the demoscene has done similar in around 1/10th that. All formant synths, of course, not the sample-based ones that you're referring to.