Knowing when to speak is actually a prediction task in itself. See eg https://arxiv.org/abs/2010.10874
Would be indeed great to get something like this integrated with whisper, LLM and TTS
Would be indeed great to get something like this integrated with whisper, LLM and TTS