Given the export restrictions this could mean they need to prioritise how to best use their limited hardware. But they could also be moving to Huawei GPUs like deepseek did and simply not have stable hardware or software for a large scale deployment yet.
This is just speculation based on the MXFP4 support on Huawei GPUs that is lacking on some nvidia GPUs.
I think the answer is that there's a tradeoff here where additional throughput for a single person can be achieved only by tying up more resources than a normal request would, even when you take into account the fact that the normal request takes longer to finish. I'm not an expert, but some of the optimizations they describe, particularly the parallel prediction stuff, sound like they could take up extra resources.
But it may well do. They mention TileRT in the announcement, so this speed comes from low level optimization for some specific GPU target.
With availability of SOTA western GPUs being scarce in China, they may well have a mishmash of different GPUs.