Qwen3 is impressive in some aspects but it thinks too much!
Qwen3-0.6b is showing even better performance than Llama 3.2 3b... but it is 6x slower.
The results are similar to Gemma3 4b, but the latter is 5x faster on Apple M3 hardware. So maybe, the utility is to run better models in cases where memory is the limiting factor, such as Nvidia GPUs?
[1] github.com/hbbio/nanoagent