I don't think it's necessarily absurd to expect accuracy from statistical methods - in many domains (including voice transcription) they blow the accuracy of traditional non-statistical approaches out of the water, and in some areas even surpass human-level accuracy.
Main thing is to measure the accuracy of the approach (regardless of whether it's traditional, statistical, or human) to determine if it's fit for purpose. In this case it sounds like the transcription shouldn't be solely relied on for high-risk decisions in its current state, but could be useful for something like searching through the reference audio if it were available.
That the issue tends to be from "pauses, background sounds or music playing" also makes me suspect a lot of the cases could be relatively low hanging fruit - check the noise gate and normalization on the microphones, or potentially have the model output a quality score for each word so that low-confidence background noise can be displayed to the end user as smaller fainter text for instance, instead of part of the conversation.