>"Smaller KV cache. Bigger savings.
Compared with the previous generation, V4.1-Flash’s KV cache needs just:
o 1/4 the HBM
o 1/8 the SSD storage
Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly."
It makes one wonder as to just how far an LLM's KV cache could theoretically be shrunk before losing significant functionality...