I tried to get Qwen3.8 27B working properly under high concurrency, and while the quality level is spectacular for the size, the performance wasn't the best, even with MTP. Unless you have a very big infrastructure, it's difficult to run a dense model concurrently with high throughput.I suppose that's why almost all large models are now MoE.
On the other hand, 1500 tok/s is an impressive speed, and that speed is very important for agent tasks, so a service like this instead of local infrastructure might make sense, although it also depends on your busines constraints.