So 45 billion parameters is what they consider their "small" model? I'm excited to see what/if their larger models will be.
So 45 billion parameters is what they consider their "small" model? I'm excited to see what/if their larger models will be.
According to Wikipedia: Rumors claim that GPT-4 has 1.76 trillion parameters, which was first estimated by the speed it was running and by George Hotz. [1]
[1] https://the-decoder.com/gpt-4-architecture-datasets-costs-an...
Also, if we have been eating up posted "benchmarks" with no way to independently validate them and watching heavily edited video presentations, why can't we trust our wonder kid?
Which also means you can fit the 8 models in a much smaller amount of memory than a 45B model. Latency will also be much smaller than a 45B model, since the next token is always only created by 2 of the 8 models (which 2 models are run is chosen by a different, even smaller/faster, model).
No, Mixture-of-Experts is not stacking finetunes of the same base model.
Made sense to mee on first sight to me, because you don't need to train stuff like syntax and grammar 8 times in 8 different ways.
Also would explain why interference of two 7B models has the cost of running a 12B model.
"Our highest-quality endpoint currently serves a prototype model, that is currently among the top serviced models available based on standard benchmarks. It masters English/French/Italian/German/Spanish and code and obtains a score of 8.6 on MT-Bench."