Thanks for pointing this out. I think we have a gap here where we may not be fully leveraging the new matmul operations available on M5+ chips, so will work that into our kernels soon and benchmark on M5 hardware.
llama.cpp had that same matmul gap (pre-fill operations) for the M5/A19 and newer silicon until that sha mentioned above.