I have tried everything (that will run on a 12GB RTX 4070) and I have yet to find anything with better accuracy than Whisper V2 Large for my dataset (discord audio from TTRPG sessions, isolated per-speaker, mostly non-American accents)
Same, for my English-only podcast
Not v3?
Voxtral to me what better
Nvidia's Nemotron subsumes their older Parakeet model now even for real time streaming.
Parakeet is way faster (on Nvidia hardware) but not quite as accurate in my experience.