Same!
There’s a reason that us humans have to use a lot of nonverbal cues in order to judge how long our responses should be, when to bail early, when someone wants to jump in briefly, beyond simply the context of the question. We even regularly alter content on the fly based on how we view the reception. Voice modes don’t have any of that context short of outright interruptions. In the meantime, some kind of response length parameter/slider would be helpful, but I think that’s a nontrivial addition in the LLM design space.
I’m curious how you were juggling this before, was it just a happy coincidence the verbosity of the replies matched your preferred pacing, or you would aggressively interrupt at times, or the model actually did a good job at conversational pacing?
Regarding length, I developed the habit of aggressively interrupting, which made voice mode basically perfect. Interrupting had to be learned because it felt very unnatural at first.
Conversely, a skill I'm currently learning is how to ask Grok to 'talk more about X' or 'can you explain that more' (I didn't need to do this prior to 2 weeks ago so I still haven't gotten good at it)