HNHacker News
TopNewBestAskShowJobs

ani17

9 karma · joined September 21, 2025

Software Developer at a computer networking company. Interested in Distributed Systems, Networking and ML.
submissionscomments
ani17··on Why do output tokens cost 5x more than input tokens?
Author here. I wanted to understand what vLLM and llama.cpp are actually doing under the hood, but the codebases are massive. So I wrote a stripped down version from scratch to see the core ideas without the production complexity.

Code: https://github.com/Anirudh171202/WhiteLotus

ani17··on LLM inference engine from scratch in C++ – why output tokens cost 5x
The blog walks through why your first token is always the slowest, why output tokens cost 5x more, and how stuff like speculative decoding and chunked prefill actually work, from the perspective of a systems engineer!
ani17··on LLM inference engine from scratch in C++ – why output tokens cost 5x
Author here. A bit more context: By day I'm a systems engineer building AI networking infrastructure. So I kept ending up in conversations where I'm not exactly able to wrap my head on the latest inference magic trick.

Like when someone mentioned vLLM's paged attention, I knew virtual memory paging, but had no idea someone had applied the same idea to KV cache allocation on GPUs.

Github link to the project: https://github.com/Anirudh171202/WhiteLotus

ani17··on Ask HN: How cam I auto-switch shared Google Meet tab?
Definitely an alternative solution. For the purpose of this script, I wouldn't prefer that though.
ani17··on How Much OpenAI Spends on Inference and Its Revenue Share with Microsoft
It's insane if the data is accurate. Only time will tell
ani17··on Taking a Look at Compression Algorithms
You forgot "Middle Out" by Pied Piper!
ani17··on Calculator Forensics (2002)
thanks for sharing!