"and $1 cost per GPU hour."
"We estimate the cost per query to be 0.36 cents."
"<2000ms latency"
"50% hardware utilization rates"
Assuming all queries take 2s this implies that ( 1 * 3600 / ( 2 / 0.5 ) ) * 0.36 = 324 GPUs would be used for each query. This seems implausibly high.