To the folks and Kitten team: I'm working on TTS as a problem statement (for an application), and what is the best model at the latency/cost inference. I'm currently settling for gemini TTS, which allows for a lot of expressiveness, but a word at 150ms starts to hurt when the content is a few sentences.
my current best approach is wrapping around gemini-flash native, and the model speaking the text i send it, which allows me end to end latency under a second.
are there other models at this or better pricing i can be looking at.