Clef is based on Qwen3.8-27B and Clef-flash is based on Qwen3.8-9B (edit: actually Qwen3.5-9B). So, similar in spirit to Kev by my understanding, but based on a newer model.
There is no official qwen 3.8 9b
From the model card:
> Clef-Flash is post-trained from Qwen/Qwen3.5-9B. See Clef for the larger variant.
16ms latency. And locally run.
Why go big when you can go small ?
To counter, most of the AI is not open. So is none of Microsoft Products. As long as they work, we keep using them.
Ai is too important and transformational to let Big Ai dominate in a closed ecosystem, thankfully the Chinese have a different mindset and approach