HNHacker News
TopNewBestAskShowJobs

vishal-padia

20 karma · joined July 2, 2023

submissionscomments
vishal-padia··on I made a kernel 2.2x faster. It made my training loop 3x slower
Quick context on what's in the post:

1. From scratch Dr. GRPO implementation in ~300 lines of PyTorch (Qwen2.5-0.5B on GSM8K, A10G). 2. Profiling deep dive on the training loop. Generate is 90% of step time. Pre-allocating the KV cache via StaticCache took GPU utilization from 26% to 86%, biggest single win in the project. 3. Wrote a fused decode-attention kernel in CuteDSL (RoPE + KV cache write + attention in one launch). Benchmarks 2.2x faster than the SDPA path it replaces at the relevant scale. 4. Plugged it into HF generate and the decode step got 3x slower. The post is mostly about why this happened, what it took to figure out, and what would actually close the gap.

Happy to answer questions.

vishal-padia··on Ask HN: What Are You Working On? (June 2025)
working on local first, bloat free MLOps tool https://github.com/Vishal-Padia/tracely

it's still wip

vishal-padia··on How I write code using Cursor
yes this!!! Whenever I write a prompt, I tend to divide it into smaller prompts, and in this process, my brain thinks of multiple ways to solve the problem. So yes, it's not limiting my thought process. I didn't notice this thing until I read this.
vishal-padia··on How I write code using Cursor
I don't agree, in the initial stages solving problems without LLMs will give a good enough knowledge about the intricacies involved and it helps develop a structured approach while solving a problem!
vishal-padia··on How I write code using Cursor
In the article, you mentioned that you've been writing code for 36 years, so don't you feel IDEs like cursor make you feel less competent? Meaning I loved the process of scratching my head over a problem and then coming to a solution but now we have AI Agents solving the problems and optimizing code which takes the fun out of it.