I know this doesn't exactly match your request, but I'm happy with my 3.6's performance. I've seen people claim they can get 120+tps on a 3090, but I'm not impatient when it comes to streaming text faster than I can read it.
I know this doesn't exactly match your request, but I'm happy with my 3.6's performance. I've seen people claim they can get 120+tps on a 3090, but I'm not impatient when it comes to streaming text faster than I can read it.
As I understand (useful) concurrency for MoE requires very large batches, where about every expert gets activated per pass.
With dense Qwen 27B on 3090/llama.cpp I get:
- no MTP: 1x42, 2x33, 3x24, 4x19 t/s
- MTP: 1x50, 2x30, 3x33, 4x30 t/sI might need to finally make the switch away from it, but it is so convenient, especially as a chat interface for system prompt experimentation.
No idea how that compares to running a larger model and context though.
I’ve sweeped concurrency across many models and many different kinds of hardware, and the only times I saw similar results to you were when I didn’t configure it correctly.