Finally we can run frontier models locally
I beg to differ on point b, but no one is delaying their next model just so they can concentrate on optimization.
It’s coming though.
One example: https://siliconangle.com/2026/07/28/ai-model-compression-sta...
Another is separating the the intelligence part of the model from the known facts part of the model (which can be better compressed)
Don't get me wrong... I think having choices is still good, I just think that having too many choices is not equally as good, or necessarily better.
Model quantization and model distillation are two techniques to reduce model size.