From using voice in the ChatGPT iOS app, I surmise that Whisper is very good at working out what you've actually said.
But it's really annoying to have to say my whole bit before getting any feedback about what it's gonna think I said. Even if it's getting it right at an impressive rate.
Given this is how OpenAI themselves use it (say your whole thing before getting feedback), I don't know that the API is set up to be able to mitigate that at all, but it would be really nice to have something closer to the responsiveness of on-device dictation with the quality of Whisper.