From 300KB to 69KB per Token: How LLM Architectures Solve the KV Cache Problemnews.future-shock.ai·157 pts·future-shock-ai·10