If you want a better experience, maybe wait for either a moe model (like 3.6 35b A3) or a model with less parameters (like 9b). Qwen has been releasing those in the past, so maybe we’ll have them for 3.8 too.
For example make an essay about something where you don't actively engage with the LLM after the initial prompt. So mostly one-shot prompts.
> It feels pretty slow on both the M5 Mac and the DGX Spark.
Probably stuck in prompt processing which is compute bound especially for iGPUs.
You've mentioned 3.5 - but it's actually the same model the only differences are training and implicit MTP support (affects prompt processing - can be disabled)
Or the "<|think_xhigh|> | <|think_low|> | <|think_off|>" tags: apart from this template detail, it is not immediately clear if reasoning_effort is deterministic (API) or is prompt engineering.