ParentFull threadyk·> a significant amount of memory goes into the KV cacheIs there a good paper (or talk) how inference looks at scale? (Kinda like ELI-using-single-gpus)View on HN