So they are comparing actual implementations with a theoretical implementation. Never mind that they got the A100 figures wrong, they are still in the 'wouldn't it be nice if we had 'x'' stage. This looks like a paper whose sole purpose is to raise funds for a research project that will probably ultimately go nowhere and they needed a reason that looks good on paper to increase their chances of getting funded. A100 can already be had for $0.87/hour so even their theoretical advantage is under significant pressure and assuming they got everything else right by the time the project has run the market will have overtaken them. This is what usually happens to CPUs that are application specific.
Pragmatically the prices are closer to $2/hr according to this recent post here on Hacker News: https://llm-utils.org/Nvidia+H100+and+A100+GPUs+-+comparing+...
Although again prices change on a daily basis on spot providers.
https://cloud.google.com/blog/products/compute/a2-vms-with-n...
That's as close as I got to verifying that price.
https://fullstackdeeplearning.com/cloud-gpus/
I feel there are more fair criticisms of that paper than its inclusion of the snapshot price of variable priced compute resource.
Of course, it's a first architectural study to illustrate the promise of the idea, lots more details to work out in a physical implementation and the final realized benefit is likely to be lower. But even a 3X is huge in this space.
> 2) Metrics: We use three performance metrics: (i) latency, i.e., end-to-end output generation time for a batch of input prompts, (ii) token throughput, i.e., tokens-per-second processed, and (iii) compute throughput, i.e., TFLOPS per GPU.
This is somewhat confusing to me because at least two of these three definitions should be essentially the same thing, but I don't think there's any way to interpret their claim of ~74 teraflops achieved other than ~211 tokens/second of throughput.
Put another way, 18 tokens per second is 2% flops utilization, which we are obviously capable of doing better than for bulk inference.
3x is not huge in this space because just using a 4090 instead of an A100 is a 5x gain.
This table was very helpful by the way, I didn't see that before. To me it clearly shows that 211 tok/s/A100 is very plausible and in fact kind of a poor showing because if you look at table D.4 and specifically the results for BS=256 PP3/TP8 they achieve ~150 tok/s/A100 on a model that's 3x larger than GPT-3.