AFAIK You can use WSE/LPU for prefill, it's just less efficient to do so.
There is also the other idea where you run your attention layer on the GPU/TPU/Trainium and the FFN on the SRAM accelerator. Because KV cache is more difficult on cerebras etc, while MOE latency is easier to deal with