This is really cool. I really enjoyed having this built into my Pixel a few years ago (super useful when you want to watch a video in public but don't have headphones). The implementation in Chrome doesn't work that well.
It would be great to support both OpenAI's whisper model and a customized vocabulary (I find that a lot of transcription errors are because the set of words I'm exposed to don't necessarily fall inside the most common 50k or 100k words).