If the resellers cache results, how it is done? I am aware it's possible with a KV cache, but those resellers do not have access to Claude's KV memory?
If they do not cache results, how does this work?
If the resellers cache results, how it is done? I am aware it's possible with a KV cache, but those resellers do not have access to Claude's KV memory?
If they do not cache results, how does this work?
My personal guess is some of those "providers" use inputs/outputs for training "open"weight models.
For example a user query is transformed to filter the request to and the answer from Claude with a local open source model.
Use case: A user asks to find a bug in code, they provide a large piece of code (hence a large amount of tokens), it is processed by the retailer with an open source LLM which is asked to extract suspicious snippets, and Claude is asked to provide an exact answer but as short as possible, then the local open source model is asked to incorporate the modifications in the original code.
The number of tokens used by Claude is much lower than if the user has asked directly to Claude. The local open source model provides the difference.