you can spawn multiple llama.cpp servers and query them simultaneously. It’s actually better this way because you get to run different models for different purposes or do sanity checks via a second model.
if I had any free VRAM at all, I would fit faster-whisper before I touch any other LLM lol