Or they simply consider inference efficiency as latency
Or they simply consider inference efficiency as latency
Could you elaborate more please? Batch inference activates pretty much all the experts since token in every sequence in a batch could hit a different expert. So at Bs=128 you’re not really getting a sparsity win.
What is you guys 70B configuration, do you guys try TP=8 for the 70B model for a fair comparison?
Which is ... a lot to say the least.
And all optimization is for latency, not throughput, because with 8 H100, you can easily hosted 4 replicas of 70B.
It lessens the number of parameters that need to be moved from memory to compute chip for each token, not from disk to memory.
It's not like there is one expert that is proficient at science, and one that is proficient in history.
For a given inference request, you're likely to activate all the experts at various points. But for each individual forward pass (e.g. each token), you are only activating a few.