Some of the models claim to be audio to audio, like one of the Gemini models. But I've tested and it does seem that you're right, it's not getting all the nuance at all
Yeah, sadly "audio to audio" seems to mean "we transcript it automatically for you internally which gets passed to the model", otherwise we'd be seeing models that are able to hear nuance in the input voice and pronunciation, which AFAIK, no model does yet.
the latest Google Translate features based on the all voice real-time 3.5 Live model is as good as I have tried in Live Modes it's not perfect but I can put it down on a table of four or five people conversing and get a reasonable amount of it translated into my earpiece.