GPT-4: 8 x 220B experts trained with different data/task distributions
twitter.com
twitter.com
Maybe this is just confirmation bias, but yeah, trying to push the model's capabilities is like working with a committee of brilliant minds chaired by an idiot.
Also, I can see why they kept this secret. Competitors just shaved months off their R&D timelines.