Could someone explain how token generation speed relates to latency for the first token to be outputted?
And if anyone have any metrics on latency on a 4090 for the 70B model, that would be very helpful.
And if anyone have any metrics on latency on a 4090 for the 70B model, that would be very helpful.