You can run the larger models just fine on a 3090. Large takes about 10G for transcribing English.
For a 1:17 file it takes:
6s for base.en, I think 2s to load the model based on the sound of my power supply.
33s for large, I think 11s of which is loading the model.
Varies a lot with how dense the audio file is, this was me giving a talk so not the fastest and quite clean audio.
While I saw near perfect or perfect performance on many things with smaller models, the largest really are better . I'll upload a gist in a but with Rap God passed through base.en and large.
edit -
Timings (explicitly marked as language en and task transcribe):
base.en => 23s
large => 2m50
Audio length 6m10
Results (nsfw, it's Rap God by Eminem): https://gist.github.com/IanCal/c3f9bcf91a79c43223ec59a56569c...
Base model does well, given that it's a rap. Large model just does incredibly, imo. Audio is very clear, but it does have music too.