Prompt Caching
docs.anthropic.com
docs.anthropic.com
https://ai.google.dev/gemini-api/docs/caching?lang=python
Costs $1 / 1M / 1h
(Yes, I realize it's probably more than 4MB, but it's still an outrageously high markup. They could do their own caching, not tell you they're doing it, and keep the difference and make even more money)
Is it a lot for caching in L1 on a chip somewhere? No that'd be wildly cheap.
Is it a lot for "caching" on a tape somewhere? Yes.
So where on this scale does keeping it quick to get to gpu memory lie?
> That's two million times more expensive than the storage cost of standard S3 (
You're not comparing to s3 at all.
Google can search the entire Internet in a fraction of a second, they can keep a million tokens within a few dozen milliseconds of a GPU for less than a dollar an hour.
If you use the Elasticache pricing, which is $0.125/gb per hour, it's still eight times more expensive. So even if a million tokens is a full gigabyte of data, it's still almost an order of magnitude more expensive than an in-memory cache adjacent to the inference boxes.
When your managed cache costs right times as much as a general purpose managed cache _in the cloud_, you've jumped the shark on pricing.
> If you use the Elasticache pricing, which is $0.125/gb per hour, it's still eight times more expensive. So even if a million tokens is a full gigabyte of data
Is it a gigabyte of data and is it fast enough?
You've guessed 4mb and 1gb. What's the actual data size here? What speed do you need to get it into the GPU ram?
The entire point here is to lower latency and costs so it has to be close and fast.
Guessing at sizes isn't helping anything here.
> A token is 32-bit integer.
No, in transformer, token is a vector, for larger models it is probably something like 6k-12k floats, assuming larger model sizes. Assume 8-bit precision, a token is more like 6-12kB, per token.
So assume 100k tokens, you will end up with 554MB for input tokens, ALONE.
Depending on your model architecture, the memory could vary, but from my observation, the runtime memory increase is at least on the same magnitude with the initial amount of memory usage upon loading the model, and this is for a moderate context length (<32k), and will grow linearly, if we don't count the n*n KV matrices.
So you are easily looking at caching 10~100GB of data, in a very hot state, and that is going to be very expensive indeed.
Size of KV cache = 2 * (num_layers) * (num_kv_heads * dim_head) * seq_length * precision
8-bit Gemma 27B KV cache = 2 * (46) * (16 * 144) * 1e6 * 1 byte ≈ 200 GB
Note that this doesn't take further optimizations into account that Google might be using.Formula: https://developer.nvidia.com/blog/mastering-llm-techniques-i...
Gemma 27B config: https://huggingface.co/google/gemma-2-27b/blob/main/config.j...
When I used the word "cost", I am referring to the price I pay, my cost, for using caching. That number is a known, published figure
My understanding is that an LLM will take in the stream of text, tokenize it (can be faster with caching, sure, but it's a minor drop in the bucket), then run a transformer on the entire sequence. You can't just cache the output of a transformer on a prefix to reduce workload.
This means that every attention layer can use previously calculated outputs for the same prompt prefix. So it only needs to calculate from scratch starting from the first unique token in the prompt sequence.
Shouldn’t be a huge deal to adjust imo
One of the bigger problems is that closed model providers don’t want to expose the embedding space and let’s users see what they have
The hard thing (I think) is what to keep in the cache and where to keep it given you are serving lots of customers and the attention calc can be a large set of numbers pretty quickly.
Plus, I'm reasonably certain they are caching the attention scores anyway.
By caching them they resume from where it left off from before thereby completely bypassing all that computation.
For large contexts this could save a ton of compute!
I think this feature and structured outputs are some of the biggest inventions in LLMs this year.
Comments suggest that caching the state of the network might also reduce processing.
I wonder if it also permits better A/B-style testing by reducing the effect of cross-domain errors. If the AI service providers made it easy to provide feedback on post-cache responses, the providers could incorporate the quality-enhancement loop accelerating time to product-market fit (at the risk of increasing dependency and reducing ability to switch).
Far less "in the realm of", "in today's fast-moving...", multifaceted, delve or other pretentious wank.
There is still some though so they obviously used the same dataset that's overweight in academic papers. Still, I'm hopeful I can finally get it to write stuff that doesn't sound like AI garbage.
Kind of weird there's no moderation API though. Will they just cut me off if my customers try to write about things they don't like?
The response you get back will have a refusal, which is pretty standard
Are contexts included in the prompt cache? Are they identified as the same or not? What happens if we approach the 10k token range? 128k? 1M?
tl;dr can cache system prompts, tools, and user messages (up to 4 total) with better returns on massive inputs such as documents.
The use case is more for client-facing applications that would hit the cache frequently rather than internal copilots.
You can put all the other source files in an initial system message and the current file in the user message. Then, if you call multiple autocompletes within 5 minutes of each other, you pay a drastically reduced price for including all of the other files in the context. Also, the latency is much reduced!
Yes, you could probably get a similar outcome by incorporating RAG, search tools etc to the autocomplete, but as a simple approach with fewer moving parts, the caching will reduce costs for this setup.
Does your approach work for you?
I'm primarily limited by how much context I need for my queries, and for the majority of the time, the context can often largely be the same across multiple queries over periods of 1-60 minutes. This is the case whether it's a codebase I'm working with or a PDF (or other form of text documentation).
Simple queries are where I expect there to be the least gain for this kind of thing.
In an ongoing conversation with a model you end up re-submitting the full test of the previous conversation - both prompts and responses - at every step. This means the cost per prompt in that conversation increases each time.
Claude prompt caching can start saving you money even with just a single user having a conversation, provided each of their replies is within five minutes of the previous reply.