Yes we do have this issue, but it's improved a bit over chatgpt due to using multiple transcribers.
The models are improving though, and they are at a very good place for English at the moment. I expect by next year we will switch over to full voice to voice models.