[see https://news.ycombinator.com/item?id=45988611 for explanation]
[see https://news.ycombinator.com/item?id=45988611 for explanation]
Great work. Learned a lot!
Where do people get the idea from that temperature affects caching in any way? Temperature is about next token prediction / output, not input.
It´s a semantics issue where the word caching is overloaded depending on context. For people that are not familiar with the inner workings of llm models, this can cause understandable confusion.
How was the term "rug" chosen, e.g. in the historical context of newspaper folds?
I'd note, when I gave the input/output screenshot to ChatGPT 5.2 it failed on it (with lots of colorful chain of thought), though Gemini got it right away.
I recently had some trouble converting a HF transformer I trained with PyTorch to Core ML. I just couldn’t get the KV cache to work, which made it unusably slow after 50 tokens…
Yes, I recently wrote https://github.com/samwho/llmwalk and had a similar experience with cache vs no cache. It’s so impactful.
I’m really glad you liked it, and seriously the resources I link at the end are fantastic.