I'm not sure that's fair, given that the distilled models are almost as good. Do you really think Deepseek's web interface is giving you access to 671b? They're going to be running distilled models there too.
Using the text "സ്മാർട്ട്", Qwen 2.5 tokenizes as 10 tokens, Llama 3 as 13, and DeepSeek V3 as 8.
Using DeepSeek's chat frontend, both DeepSeek V3 and R1 returns the following response (SSE events edited for brevity):
{"content":"സ","type":"text"},"chunk_token_usage":1
{"content":"്മ","type":"text"},"chunk_token_usage":2
{"content":"ാ","type":"text"},"chunk_token_usage":1
{"content":"ർ","type":"text"},"chunk_token_usage":1
{"content":"ട","type":"text"},"chunk_token_usage":1
{"content":"്ട","type":"text"},"chunk_token_usage":1
{"content":"്","type":"text"},"chunk_token_usage":1
which totals to 8, as expected for DeepSeek V3's tokenizer.