A niche market but I can imagine some demand there.
Biggest challenge would be Llama models.
A niche market but I can imagine some demand there.
Biggest challenge would be Llama models.
By and large, companies actually seem perfectly happy to hand pretty much all their private data over to cloud providers.
You can architect it as cheaper, low-memory GPUs, one expert submodel per GPU, transferring state over the network between the GPUs for each token. They run in parallel by overlapping API calls (and in future by other model architecture changes).
Th MoE model reduces inter-GPU communication requirements for splitting the model, in an addition to reducing GPU processing requirements, compared with a non-MoE model with the same number of weights. There are pros and cons to this splitting, but you can see the general trend.