As someone who has worked with speech to text, I can tell you from personal experience that audio seems to be broken across the ecosystem.
Bluetooth as an example is very buggy. Siri, as a voice-to-voice intelligence, fails to work most of the time for me.
I think that developers find it hard to ensure a seamless UX for anything that uses audio.