I think realtime transcription hurts the UX of polishing what's said worse. In FreeFlow the output of the transcription is fed to an LLM to polish in context of where the text is being injected. This way we can go beyond naive transcription.
FreeFlow already feels extremely fast and text being typed as I dictate is distracting especially if the polishing phase edits it.