What are the best Chinese models on HuggingFace today? Bucket by ideal RAM: <16GB, <32GB, <96, <256, 256+
Text generation, image generation, TTS, etc
What are the best Chinese models on HuggingFace today? Bucket by ideal RAM: <16GB, <32GB, <96, <256, 256+
Text generation, image generation, TTS, etc
<256: Actually, surprisingly, not a Chinese model but probably Laguna S2.1. The best Chinese model at this size is DeepSeek V4 Flash though
<96: Qwen 3.6 27B
<32: Still Qwen 3.6 27B (NVFP4)
<16: Oof, not sure. Nothing will feel great at this size TBQH without finetuning on a specific task. Pick your poison of tiny Qwen or tiny Gemma (although again Gemma is not Chinese)
- Source code (ironically on GitHub): https://github.com/modelscope/modelscope
It's not limited to Alibaba's Qwen. All of the other major Chinese ones are on there (GLM, DeepSeek, Kimi, etc.) as well as finetunes and quantizations of the non-Chinese ones.
https://huggingface.co/Tongyi-MAI/Z-Image-Turbo
Even more complex for image2video, but there’s fewer models to choose from there at least.
It would also boost research in non US jurisdiction. Who is going to be wooed by “come to our lab where you’re only allowed to work with closed models!”
All of them have quite different structures, and the reasons for choosing those structures have been clearly explained in published research papers.
The structures of the US "SOTA" models are unknown and nothing useful has been published about them, so they certainly were not a source of inspiration for China.
Big LLMs like those published by the Chinese companies must have been trained on a huge amount of text, images etc. and the training sets cannot have anything to do with the data hoarded by OpenAI and Anthropic, though they must have been gathered by the same methods, e.g. scanning the Internet and paying "pirates".
The only thing that could have been done by the Chinese companies, though for now there exists no evidence, only allegations, is that they could have used for post-training their models results of queries to US models, made by accounts which have breached the ToS, which forbid the use of the AI services by competitors.
If this really happened, this is a breach of contract, but there is no way in which one may say that the Chinese have copied anything or stolen any kind of IP and the effects of such a post-training can provide only an extremely small fraction of the information embedded in the weights of a model (though that information may be important, e.g. for ensuring that the LLM will work well in an agentic context).
I don't think any of the Chinese AI companies have stolen IP from the US AI companies (at least, I've seen no evidence of it), but the evidence of distillation is pretty apparent.
My issue with this is that if distillation occurs frequently enough, it is going to zero out most SOTA research into AI. Having cheap AI is great, but that alone won't advance the state of the art. There needs to be groups that are pushing the boundaries, and unless some kind of protection is put into place, there will be zero financial incentive to do so if anyone can come along and effectively steal your model and get financially rewarded for serving it far cheaper than the original group can, because much less R&D cost is needed. I say this as someone who sees distillation as "legal", since if you can train anything you can look at, and you can look at the output of those frontier models, then you can train on them.
Note, Laguna provided quants have some issues and they are reworking/updating them.