> It's actually how organic brains work - specialized tasks are offloaded to local cortical columns.
How are small isolated language models more similar to that than MoE in LLMs?
How are small isolated language models more similar to that than MoE in LLMs?
The original Mixtral paper [0] (in the "Routing analysis" section) found:
"surprisingly, we do not observe obvious patterns in the assignment of experts based on the topic"
A quick skim of more recent analysis on MoE shows that this hasn't changed. MoE models do appear to work, but don't appear to do what the name implies, if anything they're routing based on the structure of the text and not the semantic content (and we're still not entirely sure what they're doing).