The TTS/STT models are actually good and aggressively priced. I personally built a voice-mode ai assistant.
STT time to first token is ~300ms. ~20 second audio takes less than 1 second to be converted.
TTS time to first token is ~700ms. ~20 second of audio is generated under 2 seconds.