It is not only about computation, but also the communication bandwidth and memory, especially in training large models. The top consumer GPUs e.g. 3090/4090 still cannot beat H100 in this area even if a lot of techniques are applied.
No comments yet.