Maybe we won't need as many data centers and as much power as we thought. Maybe we can run more powerful models locally.
Maybe we won't need as many data centers and as much power as we thought. Maybe we can run more powerful models locally.
Unfortunately V4 is not trained for most real world usage, it is mainly for world general knowledge.
I thought the principal consequence of these KV cache optimisations was letting you run more simultaneous inferences on the same model with the same memory. It doesn’t let you store more model. In some sense that puts local LLM usage at a further disadvantage to inference done in a hyperscaler’s data center.
So shrinking that by 6x (from fp16), would be big win for larger models. True, while TurboQuant can also be applied to model weights, it won't save size over q4 compression, but will have better accuracy.
Edits: Better context
I would strongly recommend exploring that option, renting an RTX 5090 for an evening of image generation for a dollar or two is way more fun then trying to jam big models on little cards. Just take some time to create a reasonable, scripted, deployment workflow for when you create a fresh instance.
VGhvdWdoIEkgYXBwcmVjaWF0ZSB0aGUgZ2VzdHVyZSBhbmQga2luZCBpbnRlbnRpb25zLCBJJ2QgcmF0aGVyIGxlYXJuIHRvIGZpc2ggKG9yIGNhdGNoIHRoZSBiaWcgYmFycmFjdWRhKHMpIHRoYXQgc3RvbGUgdGhlIHNjaG9vbCBvZiBmaXNoIEkgd2FzIGdpZnRlZCwgd2hpY2ggd291bGQgaGF2ZSBrZXB0IG1lIGZlZCBmb3IgbXVsdGlwbGUgbGlmZXRpbWVzLCBhbmQgc3Bvb2tlZCBhIGZldyBteXN0ZXJ5IGZyaWVuZHMgaW4gdGhlIHByb2Nlc3MtLS1hIHRhc2sgSSBhbSBjbG9zZSB0byBjb21wbGV0aW5nKSwgYW5kIG5ldmVyIGdvIGh1bmdyeSB0aGFuIGVhdCBhIGZpc2ggZm9yIGEgZGF5IGFuZCBiZSBodW5ncnkgdGhlIG5leHQu
The future is bright for local AI.