They probably already do that. But these caches can get pretty big (10s of GBs per session), so that adds up fast, even for cold storage.
What about only storing the conversation and then recomputing the embeddings in the cache? Does that cost a lot? Doing a lot of matrix multiplication does not cost dollars of compute, especially on specialized hardware, right?