Model is suspiciously fast and has a low reported output token count (using via OpenRouter's Chat), both of which aren't representative of models from the big Chinese labs. Odd.
> aren't representative of models from the big Chinese labs
There were reports that China has let Nvidia's chips through, so this might be it. Testing both the chip and infrastructure.
They're reporting ~30tps, that's about in line with many medium sized models served by Chinese providers