And the context takes space +kv cache. KV cache drives usefulness as your context grows, it needs to pull the kv cache.
Simplified, the context has to be run on every turn, so the KV cache supplies the processed tokens, so it just needs the new inpute.
Anyway, more seriously, I hope it's obvious by now that I don't particularly care that my means of communication is so offensive to you. I think you should, like, cry a river, build a bridge, then, well... get over it, you know?
And on the, er, "topic"? A rule of thumb for me does not have to be one for you, even if it's explicitly presented as such a rule. I thought that would be obvious but, well, here we are. Anyway, how's life been treating you?
Also, in this particular thread, you started by wrongly making a correction of something that was clearly not an error, nor a wrong use of words, not even a misspelling. The problem is you started to post a correction before realizing that you didn't read right. That happens when one is more eager to boast one's own greatness than one is interested in the topic at hand. The result is that in this thread, you were clearly, as a rule of thumb, well, 100% wrong.
I have an RX 6700 XT with 12gb vram and 64gb system ram. running dense models like 27b is difficult, but i can run IQ4/IQ5 qwen 122b-a10b or 35b-a3b at ~20tok/s
Curious for any more experiences