GPT-4 Architecture
twitter.com
twitter.com
The reason companies/researchers haven't generally touched MoE for LLMs despite how good it sounds on paper is because they've typically sucked and underperformed their dense counterparts.
assuming this is all true, Did Open ai do anything differently here or is it just scale ?
I know this very recent paper shows MoE benefit far more from Instruct tuning - https://arxiv.org/abs/2305.14705
FLAN-MOE-32B comfortably surpasses FLAN-PALM-62B with a third of the compute. It goes from 25.5% to 65.4% on MMLU.
In comparison, 55.1 to 59.6% for Flan-Palm 62b. That just kind of shows the underperformance you expect from sparse models.
But from Open ai's technical report, it doesn't seem like they needed that.
The Vision component seems to be just scale. Well all of it seems to be just scale. Seems like there's plenty scale left too as far as performance gains go.
"GPT-4's details are leaked.
It is over.
Everything is here:"