Seems like a use case for AI right here; just have the user 'talk' silently, mouthing the words or in a very quiet whisper. I assume AI can read lips pretty well by now right? So turn on your video, have the AI already know how to synthesize your voice, and then just careful mouth the words into the camera. An AI model should be able to live-synthesize your voice that matches what you want to say and I think there would be minimal latency, not any worse than other video conferencing latency. And with headphones on, nobody on the plane can hear the other side, or you. And you could get a live transcript visible to you as well, so you could know if the AI made a mistake in reading your lips and you could correct it.
Might look weird but shrug.