Visualizing expert firing frequencies in Mixtral MoE
mixtral-moe-vis-d726c4a10ef5.herokuapp.com
mixtral-moe-vis-d726c4a10ef5.herokuapp.com
Edit: Oh I read it completely backwards, didn't I. Given the decisions, you can determine the topic of the input.
I don't understand why we only see 4 charts (I would expect to see 8, one for each 7B mistral tiny model)
However, this raises a question: could a slightly more complex router use output layer n-1 to choose experts for layer n+1 (vs n and n+1 today)? This way, there is more time to load the needed experts for the n+1 layer.
I'm upgrading to 96GB RAM now to run the larger models, but I do wonder whether it'll be slow when using proportionally less VRAM.
The speed is quite good even on CPU only though, I get 3.5 tokens per second with 6 cores and DDR5-6000. For comparison llama2-70B is less than 1 t/s on the same hardware in Q4. And, subjectively, Mixtral performs better.