Let's say you want to run this completely locally with Whisper and a fine-tuned LLaMA model. It there a real-time TTS that would be a good fit? The Readme only lists cloud services for TTS (text-to-speech).
We just got a PR for adding Coqui TTS which is open source – should get it merged soon :)