In this case you didn’t even get the same answer, you only happened to have one sentence in the answer match.
> Regardless of whether caching is used, the output generated will be identical. This is because only the prompt itself is cached, while the actual response is computed anew each time based on the cached prompt
That's not true at all and is exactly what prompt caching is for. For one, you can at least populate the attention KV Cache, which will scale with the prompt size. It's true that if your prompt is larger than the context size, then the prompt size no longer affects inference speed since it essentially discards the excess.
My mind immediately goes to rowhammer for some reason.
At the very least this opens up the possibility of some targeted denial of service
They probably also have cheap code or cheap models that normalize requests to increase cache hit rate.
No? Eg "how to cook pasta" is probably asked a lot.