More like `ollama launch claude --model qwen3.6:latest`
Also you need to check your context size, Ollama default to 4K if <24 Gb of VRAM and you need 64K minimum if you want claude to be able to at least lift a finger.
Also you need to check your context size, Ollama default to 4K if <24 Gb of VRAM and you need 64K minimum if you want claude to be able to at least lift a finger.
Gemma 3 27B q4:
* MLX: 16.7 t/s, 1220ms ttft
* GGUF: 16.4 t/s, 760ms ttft
Gemma 4 31B q8:
* MLX: 8.3 t/s, 25000ms ttft
* GGUF: 8.4 t/s, 1140ms ttft
Gemma 4 A4B q8:
* MLX: 52 t/s, 1790ms ttft
* GGUF: 51 t/s, 380ms ttft
All comparisons done in LM Studio, all versions of everything are the latest.