Given that Sonnet is still a popular model for coding despite the much higher cost, I expect Haiku will get traction if the quality is as good as this post claims.
Given that Sonnet is still a popular model for coding despite the much higher cost, I expect Haiku will get traction if the quality is as good as this post claims.
This could be massive.
I suppose it depends on how you are using it, but for coding isn't output cost more relevant than input - requirements in, code out ?
Depends on what you're doing, but for modifying an existing project (rather than greenfield), input tokens >> output tokens in my experience.
https://docs.claude.com/en/docs/build-with-claude/prompt-cac...
https://ai.google.dev/gemini-api/docs/caching
a simple alternative approach is to introduce hysteresis by having both a high and low context limit. if you hit the higher limit, trim to the lower. this batches together the cache misses.
if users are able to edit, remove or re-generate earlier messages, you can further improve on that by keeping track of cache prefixes and their TTLs, so rather than blindly trimming to the lower limit, you instead trim to the longest active cache prefix. only if there are none, do you trim to the lower limit.
for example if a user sends a large number of tokens, like a file, and a question, and then they change the question.
if call #1 is the file, call #2 is the file + the question, call #3 is the file + a different question, then yes.
and consider that "the file" can equally be a lengthy chat history, especially after the cache TTL has elapsed.
As far as I can tell it will indeed reuse the cache up to the point, so this works:
Prompt A + B + C - uncached
Prompt A + B + D - uses cache for A + B
Prompt A + E - uses cache for A
1) low latency desired, long user prompt 2) function runs many parallel requests, but is not fired with common prefix very often. OpenAI was very inconsistent about properly caching the prefix for use across all requests, but with Anthropic it’s very easy to pre-fire
If I'm missing something about how inference works that explains why there is still a cost for cached tokens, please let me know!
https://github.com/kvcache-ai/Mooncake/blob/main/doc/en/tran...
> Transfer Engine also leverages the NVMeof protocol to support direct data transfer from files on NVMe to DRAM/VRAM via PCIe, without going through the CPU and achieving zero-copy.
TtFT will get slower if you export kv cache to SSD.
I spend way to much time waiting for the cutting edge models to return a response. 73% on SWE Bench is plenty good enough for me.
I was hoping Anthropic would introduce something price-competitive with the cheaper models from OpenAI and Gemini, which get as low as $0.05/$0.40 (GPT-5-Nano) and $0.075/$0.30 (Gemini 2.0 Flash Lite).
This is what people mean when they say margin. When you buy a pair of shoes, the margin is price/(materials+labor), and doesn’t include the price of the factory or the store they were bought in
There are a bunch of companies who offer inference against open weight models trained by other people. They get to skip the training costs.