I haven’t tried any larger models, but a 12B model on my Ampere A5000 gets around 4-5x the aggregate throughput with concurrency. I have the maximum context configured to 32k, but my actual requests are usually around 2-4k tokens.
No idea how that compares to running a larger model and context though.