Generally infill is significantly faster than inference due to batching. Is that not the case here for some reason?
>device bytes read during the prefill window, all four drives (arm csv) 8,977 GB at 24.1 GB/s aggregate
I.e. low memory bandwidth.
They allegedly improved this by “up to 7x” with M5 but I’m not sure about the exact numbers here.