Edited to add: for agentic workflow I’m running omlx which tells me it has about a 90% cache hit rate (tradeoff is some disk and mem space) - that noticeably changes the felt speed.
I use local LLMs on my Mac Mini M4 Pro with 48G to review text messages tone, act as a text correction tool, act as a code review tool, to do code agent work, generate code snippets, etc
Gemma 4 26B A4B gives me steady 20 tps.
That is, users with M-series hardware that have less-than-max RAM share results whereas users with max RAM do not.
Speculating (not extrapolating), maybe users with machine that have max RAM are less interested in running local LLMs and are less averse to paying services for compute?
Personally, I’d love to see what output max RAM M-series Apple hardware in these threads.
I use it occasionally for classification and other tasks but I wouldn't trust those smaller models with the real work and for larger data processing it's too slow, e.g. a dataset I wanted to classify would've taken 56 days on my laptop vs just paying the cheap Luna prices to openai and getting it done in a few hours.
While I've spend a little time tuning, I'm assuming there will be deeper tuning for 3.8 that might close the gap.