I assume they didn't fix the memory bandwidth pain point though.
I'm really curious to see how things shift when the M5 Ultra with "tensor" matmul functionality in the GPU cores rolls out. This should be a multiples speed up of that platform.
Are you doing this with vLLM, or some other model-running library/setup?