"Can't" is not really correct.
Nowadays, specially with MoE models you can run parts of the model on GPU and still get some speed up.
Nowadays, specially with MoE models you can run parts of the model on GPU and still get some speed up.
Dense models are very straightforward to share/pipeline because you know all the shapes and geometry up front, that's the inference friendly option.
MoE sells a lot of HBMe3.