Fast Llama 2 on CPUs with Sparse Fine-Tuning and DeepSparse
neuralmagic.com
neuralmagic.com
DeepSparse is on GitHub: https://github.com/neuralmagic/deepsparse
Off topic: I so much appreciate everyone’s work Getting LLMs running on inexpensive hardware! I have been having crazy amounts of fun with Ollama and a wide variety of tuned models. And, so fast!
If we observe the performance comparison graph (1), 2.8->9 tok/s is achieved via quantization, but the remaining jumps from 9->16.6->24.6 tok/s are achieved from the sparse fine-tuning.
(1) https://neuralmagic.com/wp-content/uploads/2023/11/CHART-Lla...
Not all tasks require low latency.
If your usecase fits inside that 32GB (no 70B models, sadly) the price to performance of a GGUF Q4KM is really attractive on this setup.
Bigger question is, does the sparse model maintain any general knowledge?
The quantization part was interesting - they also quantized activations and have some adaptive quantization that accommodates outliers.
I tried other models and they crashed because it stored too much data in RAM