How much does that matter if it's reset at every turn?
How much does that matter if it's reset at every turn?
Your Prompt 1: Prompt Content 1 -> cache-1
LLM Response 1: <Thinking>Thinking Content 1</Thinking> Response Content 1
Your Prompt 2 (client side): prompt-1 + response-without-thinking-1 + Prompt Content 2
Your Prompt 2 (server side): cache-1 + response-without-thinking-1 + Prompt Content 2 -> cache-2
LLM Response 2: <Thinking>Thinking Content 2</Thinking> Response Content 2
Etc...
So reasoning gets dropped from context and you still get cache from the accumulating requests.Edit:
I've realised I was incorrect, the thinking doesn't get passed back and forth but the latent snapshot does which result in using memory just the same.
But that's not the full reasoning token context, just a snapshot of the latent state at the end of it, no?
Have a look at the gemini ones they're pretty small.
Looping back latent state isn't that easy. The hidden data is the entire contents of the KV cache which can be massive. I don't think any provider is trying to loop the KV cache through the client, and neural compressions of the reasoning would be lossy / an advanced technique that is still firmly in the realm of research papers, as I understand.
Gemini[0] for example passes along a snapshot of the reasoning state but it's not the equivalent to keeping all the reasoning tokens in the context.
[0] https://ai.google.dev/gemini-api/docs/thinking#signatures
Edit: Apparently it does take the same space in the LLM latent space so I was wrong.
I'm not sure if that's right. I only just recently learned that the "snapshots" (aka cached tokens) necessarily contain the entire context history, not just an image of a "state as of the final token" (well, it is the state as of the final token, but that state contains the whole history). So I'm not confident of my grasp of the structures here.