HNHacker News
TopNewBestAskShowJobs

gauravapiscean

57 karma · joined January 17, 2022

submissionscomments
gauravapiscean··on Decode-Latency Feedback Prefill: A Model-Free Controller
Author here. We tested a simple controller designed to reduce pauses while an LLM generates a response. In one setup, it reduced P99 inter-token latency by 27.7%. But it did not work on larger models or multiple GPUs. We traced the failure to the timing signal used by the controller. We published both the positive and negative results because the failure identifies an important limitation and suggests what a better controller should measure.

Happy to answer questions about the implementation, experiments, or results.

gauravapiscean··on LRU is harder to beat than the KV-cache papers suggest
Here is the repo link: https://github.com/gauravapiscean/agentic-kv-cache
gauravapiscean··on LRU is harder to beat than the KV-cache papers suggest
Author here. Context for why I did this:

There's a growing literature arguing LRU is the wrong eviction policy for agentic LLM serving, because agent sessions idle and LRU can't distinguish a paused session from a dead one. I found the argument convincing and built a simulator to exploit it. Three separate mechanisms, all lost to plain radix-leaf LRU.

The reason turned out to be more useful than the policy. When I measured — policy-independently — where recompute actually comes from on 393 real Claude Code sessions, requests arriving after a gap longer than the 5-minute provider TTL account for 17.5% of it. Requests arriving within 10 seconds account for 33.1%. The dominant waste is tight tool loops whose 88k-token working sets exceed cache capacity, not sessions idling past a TTL. That's a capacity problem, and liveness prediction can't touch it.