M2 Max processor.
I saw 60+ tok/s on short conversations, but it degraded to 30 tok/s as the conversation got longer.
Do you know what actually accounts for this slowdown? I don’t believe it was thermal throttling.
I seriously doubt it's the throughput of memory during inference that's the bottleneck here.
(A memory-bound workload like token gen wouldn't usually run into the CPU's thermal or power limits, so there would be little or no gain from offloading work to the iGPU/NPU in that phase.)