- 389 billion parameters and 52 billion activation parameters, capable of handling up to 256K tokens.
- outperforms LLama3.1-70B and exhibits comparable performance when compared to the significantly larger LLama3.1-405B model.
For the context encode it’s always close to as fast as a model with a similar number of active params.
For running on your own the issue is going to be fitting all the params on your gpu. If you’re loading off disk anyways this will be faster but if this forces you to put stuff on disk it will be much slower.
Only when talking about how fast it can produce output. From a capability point of view it makes sense to compare the larger number of parameters. I suppose there's also a "total storage" comparison too, since didn't they say this is 8bit model weights, where llama is 16bit?