HNHacker News
TopNewBestAskShowJobs

jakestevens2

20 karma · joined October 5, 2019

submissionscomments
jakestevens2··on Launch HN: IonRouter (YC W26) – High-throughput, low-cost inference
Nice! But that doesn’t answer the question. Do these optimizations don’t scale to multi-device workloads or not?
jakestevens2··on Launch HN: IonRouter (YC W26) – High-throughput, low-cost inference
Since you're using GH200s for these optimizations you're restricted to single device workloads (since GH series are SOC architecture). Kimi K2 (and many other large MoE models) requires multiple devices. Does that mean you can't scale these optimizations to multi-device workloads?
jakestevens2··on Show HN: Luminal – Open-source, search-based GPU compiler
See my other comments about static profiling of kernels. There are ways of improving the search that keep runtime at the heart of it.
jakestevens2··on Show HN: Luminal – Open-source, search-based GPU compiler
See my comment on a deeper thread about this. Eventually we will implement static profiling for common kernels so the search doesn't actually have to manually run all of them; many will have a known runtime that we can tie to them.
jakestevens2··on Show HN: Luminal – Open-source, search-based GPU compiler
Not today but we will implement memoization of kernels for each hardware backend, yes.
jakestevens2··on Sequoia backs Zed
Met the CEO of Zed. Very humble and deeply technical. Glad to see they're doing well!
jakestevens2··on Show HN: Luminal – Open-source, search-based GPU compiler
You can also set a time budget for how long you'd like the search to run for to avoid wasting time on diminishing returns.
jakestevens2··on Show HN: Luminal – Open-source, search-based GPU compiler
That depends on the model architecture and how it was written since that informs the size of the search space.

The typical range is 10 mins to 10 hours. It won't be fast but you only have to do it once and then those optimizations are set for every forward pass.

jakestevens2··on Show HN: Luminal – Open-source, search-based GPU compiler
Your description is exactly right. We create a search space of all possible kernels and find the best ones based on runtime. The best heuristic is no heuristic.

This obviously creates a combinatorial problem that we mitigate with smarter search.

The kernels are run on the computer the compiler is running on. Since runtime is our gold standard it will search for the best configuration for your hardware target. As long as the setup is mostly the same, the optimizations should carry over, yes.