I think one of the hardest things about voice AI is being able to gracefully modulate between styles of input/output delay. In actual technical conversations sometimes it is appropriate to patiently wait for someone to ask a very difficult question with delays in speech, and sometimes it is appropriate to interrupt regularly and ask brief pointed questions that require one-word answers, and everywhere in between. I'm really looking forward to having that kind of interchange with an LLM.