Not a LLM-expert, but here are three theories in descending order:
1. Quantization (e.g. fewer bits per weight)
2. Optimization to dedicated hardware
3. (Speculative) pruning of parameters to get comparable performance with a smaller model
No comments yet.