Would it be possible to reduce lag by streaming groups of ~6 tokens at a time to the TTS as they're generated, instead of waiting for the full LLM response before beginning to speak it?
- better detection of when speech ends (currently basic adaptive threshold)
- use small LLM for quick response with something generic while big LLM computes
- TTS streaming in chunks or sentences
One of the better OSS versions of such chatbot I think is https://github.com/yacineMTB/talk. Though probably many other similar projects also exist by now.
Can't wait for poorly implemented chat apps to always start a response with "That's a great question!"