For training, yeah. But at least in some cases, inference can be optimized to use less memory.
For their large Whisper model, OpenAI says they need 10GB of VRAM: https://github.com/openai/whisper#available-models-and-langu...
My DirectCompute port of that uses slightly over 4GB of VRAM, see the last column in that table: https://github.com/Const-me/Whisper/blob/master/SampleClips/... I haven’t actually optimized memory usage I only did the DirectCompute port, the memory savings were achieved by Georgi Gerganov in whisper.cpp project, which I ported to Windows.