If you pre-process a text file, you cannot insert it into the context because even if you numbered the tokens within a file properly, it is numbered wrong in the global context of the request. So you have cached data but the request structure does not allow you to insert it.
I think one way out of this would be to split a token into its request and content components. Then you might have to recalculate the request level information but never the content level information.
Edit: Now that I think about, why even cache the RoPE'd values to begin with? If you only need the RoPE'd data during token generation, then you could just RoPE on the fly instead of baking it into the KV cache, meaning you unlock block level suffix and infix caching, not just prefix caching.
Edit 2: In case caching the RoPE'd values is necessary, it might still be possible to apply the inverse RoPE for recalculation purposes.
Edit 3: By reserving a fixed block of tokens for a summary at the front right after the system prompt, you could now update that summary for the cost of prompt processing the summary without worrying that you modified something at the front.
The issue is that when you have messages 1, 2, 3, 4, 5. And you want to update messages 1,2,3, because they are in the reserved chunk, then 4 and 5 already contain some baked in information from 1,2,3. This means you cannot modify 1, 2, 3 arbitrarily without changing the architecture, but inversely speaking, if all you are doing is generating a summary that is strictly meant to reflect the content of the messages 1, 2, 3, then replacing them with a less information dense version that still retains the same meaning for 4, 5 means you could implement this directly in the inference engine without modifying the model at all. This would just be a special form of sliding window attention where the prefix is updated and causes partial preprocessing.