> 2) Metrics: We use three performance metrics: (i) latency, i.e., end-to-end output generation time for a batch of input prompts, (ii) token throughput, i.e., tokens-per-second processed, and (iii) compute throughput, i.e., TFLOPS per GPU.
This is somewhat confusing to me because at least two of these three definitions should be essentially the same thing, but I don't think there's any way to interpret their claim of ~74 teraflops achieved other than ~211 tokens/second of throughput.
Put another way, 18 tokens per second is 2% flops utilization, which we are obviously capable of doing better than for bulk inference.
3x is not huge in this space because just using a 4090 instead of an A100 is a 5x gain.