HNHacker News
TopNewBestAskShowJobs

gkapur

118 karma · joined July 7, 2016

I work in venture capital.
submissionscomments
gkapur··on [dead]
Currently, the most efficient kernels come from vendor libraries like cuBLAS or hand-optimized libraries like FlashAttention. Compiler-generated kernels (TVM, Hidet, Inductor, etc.) can generate more efficient kernels for some specific operations, but fall short of cuBLAS and FA on common general matrix multiplication and attention shapes. We demonstrated that a compiler can match or even outperform vendor libraries. On Gemma 4 12B, a small open-weight model, running on RTX 5090, compiler-generated kernels achieved up to 1.6× speedup over cuBLASLt, and a 1.30× geometric-mean speedup over PyTorch. Primarily, two optimizations allowed the Emmy compiler to outperform cuBLAS on common GEMM and FA shapes on RTX 5090 and RTX 4090: TMA transport (Blackwell-only) and leveraging full FP16 tensor cores with FP32 shadow accumulation registers:

TMA Transport for Matmul and Flash Kernels. cuBLAS and Flash Attention kernels still use the same cp.async transport on consumer Blackwell dies (in fact, most cuBLAS GEMM kernels on consumer Blackwell, including the FP16 tensorop path, are forward-ported Ampere-era cutlass_80_* kernels). Swapping cp.async with TMA allows us to reduce the number of instructions kernels need to issue, and the TMA’s swizzle drops shared-memory bank conflicts for free.

Hybrid FP16/FP32 Accumulation. The default matmul path on the production stack uses FP16 tensor cores with FP32 accumulation. However, on consumer dies, the FP16-input/FP32-accumulate HMMA runs at exactly half the rate of FP16-input/FP16-accumulate. To work around this, I use the fast atom mma_m16n8k16_f16_f16, but keep accuracy in check by promoting the FP16 partials into the FP32 registers and thus doing global accumulation accurately in FP32. So you get the FP16 tensor-core speed with an FP32 accumulation instead of paying the FP32-accumulate tax on every single mma. We measured the error of C = A@B using FP16 inputs drawn from N(0,1), comparing each accumulation strategy against an FP64 reference over the identical FP16-rounded operands. All configurations land within one standard error of each other, so the hybrid accumulation does not degrade task quality, matching the kernel-level error analysis.

The compiler approach allowed us to introduce both of these optimizations into all GEMM and Attention kernels used in the model rather than requiring hand optimization.

Repo: https://github.com/cloudrift-ai/emmy

gkapur··on [dead]
Currently, the most efficient kernels come from vendor libraries like cuBLAS or hand-optimized libraries like FlashAttention. Compiler-generated kernels (TVM, Hidet, Inductor, etc.) can generate more efficient kernels for some specific operations, but fall short of cuBLAS and FA on common GEMM and attention shapes.

We were able to demonstrate that the compiler approach can match or even outperform vendor libraries. On Gemma 4 12B, a small open-weight model, running on RTX 5090, compiler-generated kernels achieved up to 1.6× speedup over cuBLASLt, and a 1.30× geometric-mean speedup over PyTorch.

Primarily, two optimizations allowed the Emmy compiler to outperform cuBLAS on common GEMM and FA shapes on RTX 5090 and RTX 4090: TMA transport (Blackwell-only) and leveraging full FP16 tensor cores with FP32 shadow accumulation registers:

TMA Transport for Matmul and Flash Kernels. cuBLAS and Flash Attention kernels still use the same cp.async transport on consumer Blackwell dies (in fact, most cuBLAS GEMM kernels on consumer Blackwell, including the FP16 tensorop path, are forward-ported Ampere-era cutlass_80_* kernels). Swapping cp.async with TMA allows us to reduce the number of instructions kernels need to issue, and the TMA’s swizzle drops shared-memory bank conflicts for free.

Hybrid FP16/FP32 Accumulation. The default matmul path on the production stack uses FP16 tensor cores with FP32 accumulation. However, on consumer dies, the FP16-input/FP32-accumulate HMMA runs at exactly half the rate of FP16-input/FP16-accumulate. To work around this, I use the fast atom mma_m16n8k16_f16_f16, but keep accuracy in check by promoting the FP16 partials into the FP32 registers and thus doing global accumulation accurately in FP32. So you get the FP16 tensor-core speed with an FP32 accumulation instead of paying the FP32-accumulate tax on every single mma. We measured the error of C = A@B using FP16 inputs drawn from N(0,1), comparing each accumulation strategy against an FP64 reference over the identical FP16-rounded operands. All configurations land within one standard error of each other, so the hybrid accumulation does not degrade task quality, matching the kernel-level error analysis. The compiler approach allowed us to introduce both of these optimizations into all GEMM and Attention kernels used in the model.

Repo: https://github.com/cloudrift-ai/emmy

gkapur··on Exploring FlashAttention-3/4 optimizations on RTX GPUs
I was curious whether any of the FA-3/4 optimizations transfer to RTX GPUs. vLLM/SGLang attention falls back to FA-2 on consumer cards (FA-3 and FA-4 are datacenter-only), so I wanted to know if there's any performance left on the table, and I rebuilt the attention kernels from scratch.

The kernel reaches parity with FA-2 (206us on RTX5090 with batch=1, heads=8, seq_len=4096, head_dim=64), but unfortunately, FA-3/4 optimizations are either not applicable or not helpful on consumer cards. It looks like FA-2 is the ceiling.

In summary: - Faster tensor-core instructions (WGMMA) are the main lever behind FA-3, but they are not available on RTX GPUs. - TMA (tensor memory accelerator) is available on sm_120 (RTX 5XXX). It helps on paper (LSU drops), but the transport isn't the bottleneck, so the final number barely moves. - Warp specialization is also available. However, it is mainly a scheduling optimization, i.e., it helps eliminate pipeline bubbles and better utilize tensor cores, but without asynchronous tensor core instructions. The result is negative: 213 vs 206 us. - In FA-4, they also simulated exp using FMA instructions because the tensor cores on the B200 are so fast that the whole pipeline became SFU-bound (special functions unit). RTX 5090 is tensor-core bound, so no point in this optimization either. In fact, even a conventional optimization of using faster exp2f instead of expf for softmax doesn't move the number.

I have tried a handful of other optimizations that could potentially work on consumer silicon, such as a deeper pipeline and register ping-pong. No luck. Since the whole pipeline is tensor-pipe-bound, I believe the FA-2 is the ceiling, and that all meaningful levers will require sacrificing some accuracy to leverage faster, lower-precision tensor cores.

Note that this is an exploration of the regular attention that dominates the prefill- and compute-bound regimes. Decoding against a large KV cache is a different, memory-bound story where split-KV/Flash-Decoding matters more than any of the above.

Github: https://github.com/cloudrift-ai/emmy

gkapur··on Ollama: All Aboard Open Models
Peter Fenton invested in the company when it was called Infra.App originally (I think it's still online: https://infra.app/). It was an access management product, then it became desktop Kubernetes product, both they pivoted out of. I imagine he invested based on the technical acumen of the founders and then worked with them to pivot when Ollama started working out. Almost all of the other investors invested in the founders based on that original idea space (which wasn't a great one but was quite adjacent.) I think that includes 8VC, YC, Pace Capital, GTMFund, who all followed on here.

Since it pivoted, I imagine the story they are telling is the usage. But you can add to that Peter Fenton has had incredible success in commercial open source (Docker-ish, Elastic, JBoss, CockroachDB, TimeScaleDB, and if you include Benchmark there is Elastic, Confluent, etc.) so a bunch of these people are probably trying to ride the Peter Fenton/Benchmark train. My guess is that's what Theory is doing.

Investors are not free but they are very hype driven. There are < 10 investors who are truly experts in open source in my opinion in the market (and maybe < 5.) It's a really wacky cadence and I would argue there are different types of OSS businesses, which makes it even more difficult to be an expert. It's something I struggle with personally investing in OSS -- I think I understand the likely motion for an OSS startup and then > 50% of the time I am wrong.

gkapur··on Inkling: Our Open-Weights Model
The story of Reflection AI is supposedly that the company was faffing and failing at winning in the coding agent space, but was introduced to Jenson, who suggested they build an open-weight model and said he would fund it. That turned into a $2 billion financing with NVIDIA doing roughly $500 million and was a complete pivot.

I think the bet would have to be that a US Open Weight company either: 1. Gets a lot of money from Jenson who views them as a counterbalance to the big labs in his ecosystem and a way to generate leverage (the same way he is positioning neoclouds-- it also could be synergistic with neoclouds who could offer the model serving endpoints) 2. Can fast follow the same way Mistral does (which, honestly, seems like just distilling the Chinese model, which distills the US lab but is pretty innovative on a whole lot of architecture both in training and serving land.) 3. AND figure out some (maybe not super lucrative but lucrative enough) sort of business model, as well. There are lots of possible business models, so I will be curious how this whole space evolves.

gkapur··on Inkling: Our Open-Weights Model
If they have a really seamless fine-tuning experience and maybe can help you extract the data you need to FT (which is one of the big challenges in actually getting fine-tuning democratized), maybe you would use it because "Tinker" defaults to it.

The model could also be more flexible for non-coding use-cases (they show the results for reasoning being strong) so maybe the argument is to use it for non-coding use-cases to drive relatively deterministic conclusions for non-coding agents (they have also done some determinism work on kernels, which could be useful in pulling on that thread of deterministic models that are fine-tuned for everything that is not writing code.)

That said, I'm not sure how much all the work they have done actually synergizes or if the market size (at least in the short to medium term) is big enough for a huge outcome from the company's current valuation with those bets as the enterprise agent estate is taking a while to evolve. Hence companies like Anthropic and OpenAI are throwing tons of consulting money at the problem.

gkapur··on Inkling: Our Open-Weights Model
It could be but there are a host of companies going after open weights models: Arcee, Reflection, Llama (TBD on Meta's focus on closed-source versus open-source), etc.

That said, the fine-tuning API + open weight model at least is a semblance of a viable business that could work so I will be curious about it. I'm not sure the synergy is fully there (why is someone with an open weights model privelaged to fine-tune it better if it's just QLora or Lora) but let's see!

gkapur··on Tutorial: Algebraic Foundations Powering FlashAttention
I'm writing a short series of tutorials on FlashAttention: from theory to efficient CUDA kernels.

Part 1 is the theoretical foundation. It walks through a modern algebraic formalism showing that FlashAttention is an associative operation, which lets you treat it as a regular reduction on the GPU and apply all the same scheduling optimizations. Some recent MLSys and CVPR (ELSA) papers lean on this framing, and I find it much more powerful than the original.

This framework is particularly useful in ML compilers, where you should implement general optimizations applicable to many operations rather than writing specialized kernels. This article shows that attention actually belongs to a large family of "secretly-associative" operations, walks through a handful of examples, and links a few concepts from abstract algebra that let you identify whether an operation is secretly associative.

Overview:

- Safe softmax, Welford's variance, and FlashAttention are the same secretly-associative operation

- The twisted monoid (transport of structure), why the max-rescale coupling doesn't break associativity

- The qk_scale = log2(e)/√D like in FA-2 derived from scratch

- Numerical analysis: overflow bounds, error limits, and why tiling never amplifies error

- Bird's 3rd Homomorphism Theorem as a test for whether any loop is secretly associativ

gkapur··on Orthrus-Qwen3: up to 7.8×tokens/forward on Qwen3, identical output distribution
On the limitation side:

Do you think this would scale to larger transformer models with more parameters per layer?

How would this work with MOE models or sparse models?

gkapur··on AdapTive-LeArning Speculator System (ATLAS): Faster LLM inference
Adding to the prior comments as my intuition matched yours, there’s a nice Reddit thread that gives some context into how it can be faster even if you require exact matches: https://www.reddit.com/r/LocalLLaMA/s/ARxHLqRjdM

The TLDR/key (from my understanding) is that verifying N tokens can be faster than generating N tokens.

gkapur··on Fivetran Acquires Dbt Competitor Tobiko Data
Congratulations to their team. SQLGlot is a really powerful tool that a lot of companies use so a huge contribution to the OSS community so hopefully it continues to be supported and gets better and better!
gkapur··on Darklang Goes Open Source
There was also Wing cloud (fka Monada) and there’s Mojo by Modular (https://www.modular.com/mojo.)

Feels like two types of companies raised money: - Companies trying to couple the cloud with a programming language. - More recently, companies trying to couple GPUs with a programming language/alternative to CUDA.

Will be curious how this generation goes.

gkapur··on Show HN: Agno – A full-stack framework for building Multi-Agent Systems
If you are running things locally (I would think especially on the edge, whether on not the LLM is local or in the cloud) this would matter. Or if you are running some sort of agent orchestration where the output of LLMs is streaming it could possibly matter?
gkapur··on FTC takes action against Uber for deceptive billing and cancellation practices
I’m convinced I get more “deals” (temporary discounts) from Uber without Uber One/after canceling it, which offsets the benefits from Uber One.

I don’t see those deals on Uber Eats so it feels like the real value of Uber One is for heavy Uber Eats users.

PS. Worth going through the cancellation flow when you are up for renewal as they will probably offer you 50% off Uber One.

gkapur··on Comparing Auth from Supabase, Firebase, Auth.js, Ory, Clerk and Others
Today there are so many other solutions: Stytch, Descope, PropelAuth (For B2B companies), and others.

VCs went a bit ham on this category when Auth0 got bought. I sense that the general thought process was: Auth0 multi-billion dollar company -> Auth0 will become crappy under Okta > replace Auth0 and build a bit company.

gkapur··on Dbt – Incremental but Incomplete
Basically people are constantly calculating metrics based on existing tables. Think something as simple as a moving average or the sum of two separate columns in a table. Once upon a time you would set up a cronjob and populate these every day as a SQL query in some python or Perl script.

Dbt introduced a language for managing these “metrics” at scale including the ability to use variables and more complex templates (Jinja.)

Then you do dbt run (https://docs.getdbt.com/reference/commands/run) and kapow the metric is populated in your database.

More broadly dbt did two other things: 1. It pushed the paradigm from ETL to ELT (so stick all the data in your warehouse and then transform it rather than transform it at extraction time.) 2. It created the concept of an “analytics engineer” (previously know as guy who knows SQL or business analyst.)

gkapur··on Command AI Bought by Amplitude
Thanks for the transparency and thoughts!
gkapur··on Command AI Bought by Amplitude
What’s interesting is how much it contrasts with TechCrunch’s story: ‘Most of Command AI’s 30-person, San Francisco-based team will be joining Amplitude. Command AI’s co-founder and CEO James Evans wouldn’t reveal the terms of the deal, but said candidly that an acquisition wasn’t something he’d been planning on. “Our growth was great and we had plenty of runway,” Evans told TechCrunch. “We weren’t out shopping ourselves or anything. But when Amplitude reached out a little while ago — this summer — we got really excited about the combination and became convinced that we could grow faster and reach more users together.”’
gkapur··on Command AI Bought by Amplitude
Interestingly, according to Axios, the price was pretty limited: "Amplitude (Nasdaq: AMPL) acquired CommandAI, an SF-based software user experience startup, for $20m (net of cash). CommandAI (fka CommandBar) had raised around $23m from Insight Partners, Itai Tsiddon, Thrive Capital, and BoxGroup."

I would be curious to learn more about the rationale to sell the business as I understand the strategic value to Amplitude. Interestingly, these next-generation digital adoption platforms have generally been pretty challenged.

gkapur··on Show HN: Vortex – a high-performance columnar file format
Not an expert in the space at all and it does seem like people are exploring new file and table formats so that is really cool!

How does this compare to Lance (https://lancedb.github.io/lance/)?

What do you think the key applied use case for Vortex is?

gkapur··on Response to DHH
Been watching this episode unfold on Twitter and has read about Matt’s domain hijacking of thesis, etc.

It seems to me like Matt is the type of person who likes to hide behind character and other ad hominem attacks rather than addressing actual issues at hand. Perhaps this is because of a psychological issue but I can’t really know. Normally I would think the community would be repulsed by this and would find an alternative.

What is remarkable is that there has not been enough community animus to fork relative to Terraform and OpenTofu. It shows the power of the underlying GPL license and “plugin” approach. Something to think about for other companies that may be thinking about relicensing versus just building a durable ecosystem around their proprietary brand and assets.

gkapur··on Milvus Lite: The Lightweight Version of Milvus
Not really. This is more like SQLite or DuckDB for vector databases (on disk.) Chroma is more like redis for vector databases (in memory.)

We have seen similar products in the olap space, as well, ie. Clickhouse local.

gkapur··on Bento: Open-source fork of the project formerly known as Benthos
This whole thing comes off as tone-deaf and deceptive even to me (who is all for COSS monetizing.) Warpstream was sponsoring Benthos, it sounds like they didn't get a great heads up of this happening, which makes the project owner sound self-serving. Then you renamed the repo and relicensed some connectors all in one go without giving anyone from the community a chance to opine or think about how this affects them.

Finally, Redpanda did some partnerships with vendors nobody cares about whose businesses are at risk to show how you are opening up the ecosystem.

It actually comes off as somewhat malicious and Ashley's note where he notes he didn't read the article also comes off as not caring about developers (even insofar as he has facts wrong -- if the plug APIs remain compatible this creates more choice for users.)

gkapur··on The end of Airplane.dev
I empathize for the co-founder who was CTO and became CEO. I imagine some of the challenges come from the fact that there was a big chunk of equity owned by the original founding CEO. As the remaining co-founder, I can imagine feeling like I was climbing a grueling uphill battle to just get back to the most recent $300 million (!) valuation and a large amount of the upside of that battle was going to someone who abandoned me/the team. That said, I'm not privy to any equity adjustment that happened after the original founding CEO left.

That said, the after-the-fact communication feels lacking. It sounds like your CEO could have at least explained things in more detail after the fact (ie. Why he felt a pivot would be necessary to get to where the business needed to go.)

Such is life, I guess. I'm wishing you and the team the best for the future!

gkapur··on Comparing Postgres Managed Services: AWS, Azure, GCP and Supabase
Do what it’s worth supabase definitely feels slow to me. Neon, in contrast, feels lightning fast for my workloads.
gkapur··on Phidata: Build AI Assistants using function calling
Very cool! Function calling seems to be a new paradigm that is really taking off.

Does anyone know how the LLM vendors are actually implementing function calling? Is it just a thoughtful prompt and a loop where they parse outputs and check if it corresponds to the arguments of the function? Or something else?

gkapur··on Show HN: Epsio – Incremental views for your existing database
Cool product! A few questions mostly out of curiosity:

- What underlying data flow technology is this based on? Is it Timely dataflow a la Materialize, something else a la RisingWave, etc.? - What can the materialized views not have? Ie. Does it support window functions? Unions? User-defined functions? - Finally I’m curious about how you think of the landscape versus RisingWave, Readyset, Materialize, etc.!

gkapur··on OpenTerraform – an MPL fork of Terraform after HashiCorp's license change
> Do they not want to anyone offering anything that could compete with any of Hashi's products to be able to use terraform binary at all?

I think they want to make it difficult for SaaS products that embed terraform (or an altered version of terraform I would assume where they have built APIs) in their product. That said, customers can install terraform so you can still hypothetically have an "integration" with Terraform where you call the raw binary (like Atlantis). So basically, they want terraform cloud to have a "privileged" user experience is my guess.

gkapur··on OpenTerraform – an MPL fork of Terraform after HashiCorp's license change
Don’t think it’s targeted at Vault. Vault has APIs so in the enterprise, you can self host and a vendor can manage vault through the APIs. Terraform specifically is built CLI only in the OSS so you have to repackage the underlying binary and hence are subject to the license as a vendor.
gkapur··on Show HN: DevPod – Codespaces but Open Source, Client-Only, and Unopinionated
Congratulations on getting the product out! Seeing innovation in local development, which seems to satisfy a lot of use cases with MacBook compute and memory increasing all the time, with the flexibility of using the cloud will be super helpful to some companies looking to save on cost!
Page 1 of 2Next →