how're you handling latency on turn overlaps : buffered stream with early intent cutoff or full duplex with partial decoding?
Whisper can transcribe in <100ms. We then wait for the turn detection model, LLM, and tts to trigger a streamed response back to eh client.