One idea I was tossing around was streaming transcription + batch re-transcription:
- Use streaming transcription, which works most of the time (for example, I've found the Web Speech API pretty good, as well as moonshine)
- If the streaming transcription was poor, select the bad part and re-transcribe with a more accurate batch transcription model.