M4 Max is typically better than M5 Pro for inference IIRC.
Time to 1st token is faster on the M5 because of HW accelerators helping the prompt interpretation (and it is CPU-bound).
Token generation after that is GPU-bound and will profit from the higher bandwidth of the M4 Max.