> Then the part that matters: where the KV lives
When your abstract was clearly generated by an LLM and not curated to at least make it sound human, it does not make me want to read your paper.
When your abstract was clearly generated by an LLM and not curated to at least make it sound human, it does not make me want to read your paper.
KV caching is a super interesting engineering space, especially when you’re talking about local models where compute and memory bandwidth are highly constrained and you’re trying to trim fractions of a second everywhere you can by flipping between different ICL prefixes. But selling caches for specific documents just makes no sense at all.