This thing fly on Macbook M4 Max 128GB at over 100t/s, for small contexts, over 20t/s for large contexts. MLX 4bit quant.
It would be nice having a fast local model that is good at using tools
Have you tried using them with something like Claude code or aider?