Actually I was just checking, and Bark isn't that close to maxing out GPU utilization. Running two instances on a 3090 seems like a throughput increase and the models fit. Update: And getting weird CUDA issues. Hmn...
Just to add a datapoint: the main audioLM based models (not the BERT embedding part) fully utilize an RTX 2080 Ti.
Must be some low hanging fruit to optimize in Bark. It would be somewhat close to realtime if it was close to 100% and scaled linearly.