Totally agree — the “demo vs real world” gap is always the messy edge cases: accents, crosstalk, domain terms, and people talking like… people.
Did you end up adding any guardrails (confidence thresholds, “please repeat,” glossary/term injection, or human fallback)? Also curious: were failures mostly ASR or translation/context?