However, code changes are necessary to achieve that, although they won't be crazy complex.
> However, code changes are necessary to achieve that, although they won't be crazy complex.
This is technically true. It will be very slow though.
However, give it 6 months and I think we might see an order of magnitude increase in speed on CPUs. This will still be too slow to be very useful though.
If you have a guess what the model will output, then you can verify that your guess is correct very cheaply, since you can do it in parallel.
That means there is the possibility to have a highly quantized small model in RAM, and then use the big model only from time to time. You might be able to get a 10x speedup this way if your small model agrees 90% of the time.
The potential for speedup according to their paper is closer to 2x than 10x however.
Erm, for inference that is. Training is definitely out of question for individuals I believe (unless you use much smaller models?).