They're not doing autoregression, so all the outputs are computed in one big forward pass. Very cheap.
I think OP is confused about "others" vs "them".
they’re talking about two totally different things, right?
The output tokens are just responses to your inputed questions and their probability. So relatively few output tokens. No unstructured text back in the response.
it's our output tokens that are free (under the system one / jev column)