Is that separately comparing the time it takes to preprocess the input prompt (prompt_length / pp_token_rate = time_to_first_token) and then the token generation rate is the time for each successive token?
I also see something about bs batch size. Is batching relevant for a locally run model? (Usually you only have one prompt at a time, right?)