HNHacker News
TopNewBestAskShowJobs

gajjanag

433 karma · joined July 25, 2015

Software engineer with a keen interest in high performance computing. In the past, PhD from MIT in EECS in a mix of applied mathematics and computer science.

Research interests: a broad range of applied mathematics topics, generally in the neighborhood of information theory and probability.

Other interests: mathematics in general, computer systems (especially security), FOSS projects, reading, bicycling, and cooking.

Website: gajjanag.github.io

[ my public key: https://keybase.io/gajjanag; my proof: https://keybase.io/gajjanag/sigs/-vv1qyl9QGR_46BS-LPAm3wBVPYna8zjMyUzDBWVKWg ]

submissionscomments
gajjanag··on We must pace the frontier
> The entire essay is about stopping development until it can be made more safe, with specific ideas on how to do that.

No, the essay does not talk about "stopping development". That's a very clever sleight of hand made in the essay to get readers to draw this false conclusion.

It just talks about building AI at a "balanced rate". Untangling the corporate speak, this amounts to essentially "go full steam ahead, but have more eval oversight before release".

gajjanag··on Everyone should slow down AI development except for me
> The idea that open-weight models are more efficient seems unfounded.

https://artificialanalysis.ai/ intelligence vs cost per task disagrees with this statement.

gajjanag··on How I find problems to solve as a staff engineer
> You will typically find your staff and higher level engineers rarely are actually writing code themselves.

Many so called staff+ engineers are basically glorified Google Docs engineers, or management who want to call themselves technical. It seems to be fashionable these days to call oneself "technical" without actually knowing anything of the problem domain.

They basically take all the "credit" for any innovation out of the team, and scapegoat other teams when project deliverables are not met.

gajjanag··on Amazon workers under pressure to up their AI usage are making up tasks
> Communication matters most when you're dealing with cross-org concerns and those that master it are usually the more friendly and pleasant ones.

I don't agree with the second one, but agree with the first.

Throughout my corporate career so far, I have found plenty of hot air/pretty picture slide decks that exist solely for ladder climbers to climb. Said ladder climbers are usually all smiles in public and "friendly", but you have to watch out for knives behind your back.

gajjanag··on TurboQuant: A first-principles walkthrough
> Thanks for that! It is worth noting that taking advantage of the post-rotation distribution

I again feel this claim is too strong. Rotations have been used in information theory/wireless communications for decades at this point, with appropriate scaling done at channel inputs/outputs to hit channel capacity. The signals then pass through the appropriate codebooks that take advantage of the post-rotated+whitened signal.

Our cellphones today are powered by such technology.

I agree with your claim when restricted to deep learning. But I do not agree with the broad characterization that taking advantage of post-rotation distributions was only first done in your work.

gajjanag··on TurboQuant: A first-principles walkthrough
Wow, yes - you are completely correct (read through the note in detail now).

Though, as your paper also notes, the quantizer values themselves aren't fundamentally novel to either paper. Lloyd Max scalar quantizers have been studied for a very, very long time. And the specific Lloyd Max values for the Gaussian input distribution have been obtained in many papers across signal processing and information theory.

gajjanag··on TurboQuant: A first-principles walkthrough
There are also more papers on similar themes.

For example, TurboQuant makes use of QJL (quantized Johnson Lindenstrauss transformations). One of the first papers to characterize the QJL and in fact the rate distortion tradeoff for quantized matrix multiplication in general is "Optimal Quantization for Matrix Multiplication" (https://arxiv.org/abs/2410.13780) by Ordentlich and Polyanskiy.

There is also a more accessible survey paper around quantized matrix multiplication called "High-Rate Quantized Matrix Multiplication: Theory and Practice" (https://arxiv.org/abs/2601.17187), by the same authors.

TurboQuant cites none of them.

gajjanag··on The RAM shortage could last years
TurboQuant is known across the industry to not be state of the art. There are superior schemes for KV quant at every bitrate. Eg, SpectralQuant: https://github.com/Dynamis-Labs/spectralquant among many, many papers.

> Given that TurboQuant results in a 6x reduction in memory usage for KV caches

All depends on baseline. The "6x" is by stylistic comparison to a BF16 KV cache; not a state of the art 8 or 4 bit KV cache scheme.

gajjanag··on The GNU libc atanh is correctly rounded
The bigger challenge is GPU/NPU. Branches for fast vs accurate path get costlier, among other things. On CPU this is less of a cost.

Most published libm on GPU/NPU side have a few ULP of error for the perf vs accuracy tradeoff. Eg, documented explicitly in the CUDA programming guide: https://docs.nvidia.com/cuda/cuda-programming-guide/05-appen... .

Prof. Zimmermann and collaborators have a great table at https://members.loria.fr/PZimmermann/papers/accuracy.pdf (Feb 2026) comparing various libm wrt accuracy.

gajjanag··on Don't become an engineering manager
> That's why you need to put your scope

The problem is, "scope" is often equated to "how many people worked in my empire" rather than "how much business value did my work X generate".

The two things are vastly different, and I have seen the distinction/oversimplification play out over and over in my own career as well as many others around me.

As an extreme on the "individual technical expert side", there are things out there that can pretty much only be accomplished with a few people around the world who possess the dedicated expertise. These results can't be replicated by a cobbled together team of 10 or 100 people even though the latter sounds more impressive for "scope".

Some organizations do a decent job of recognizing these different "archetypes", many don't.

gajjanag··on The state of SIMD in Rust in 2025
>80%-90% or so of real life vectorization can be achieved in C or C++ just by writing code in a way that it can be autovectorized.

Yep. I was pleasantly surprised by the autovectorization quality with recent clang at work a few days ago. If you write code that the compiler can infer to be multiples of 4, 8, etc the compiler goes off and emits pretty decent NEON/AVX code. The rest as you say is handled quite well by intrinsics these days.

Autovectorization was definitely poorer 5-10 years ago on older compiler toolchains.

gajjanag··on Advice for new principal tech ICs (i.e., notes to myself)
Welcome to the brave new world these days:

1 - Very few people conduct "proper scholarship", and fail to trace ideas back to their original inception and cite them correctly. This happens time and again in deep learning, where 30+ year old ideas are claimed as "novel" over and over. Many times out of malice by the authors, sometimes out of ignorance.

2 - Peer review in many parts of the industry+research is a joke. Mostly shouldered by early graduate students who don't really know the field well and an incredibly noisy process.

3 - It is common practice now to dump out one's "kitchen sink" of ideas rather than properly refined stuff. Hence the increase in LinkedIn spam, blog spam, arXiv spam style of papers.

gajjanag··on In Defense of C++
> I don't think there are many (or any) upsides to the well documented downsides.

C++ template metaprogramming still remains extremely powerful. Projects like CUTLASS, etc could not be written to give best performance in as ergonomic a way in Rust.

There is a reason why the ML infra community mostly goes with Python-like DSL's, or template metaprogramming frameworks.

Last I checked there are no alternatives at scale for this.

gajjanag··on Defeating Nondeterminism in LLM Inference
As others have pointed out, these phenomena are well known to many folks across companies in the AI infra space. It doesn't really break new ground. This article is a good exposition of the basic strategies though.

What I would have loved is a discussion around collectives/multi-node setups. And showing how to get determinism at low performance penalty for multi-node reduction collectives.

gajjanag··on SF may soon ban natural gas in homes and businesses undergoing major renovations
I guess you have never worked with a slow induction cooktop. Literally we had to spend 15 minutes more for cooking things on induction compared with our previous apartment's gas connection.

Maybe they are better now but it is certainly not the case that all induction cooktops have these magical properties; many are cheap and skimp on something. While in the 5+ apartments I have been in gas has always delivered the same heating experience that I can rely on.

And to your point about rotis, no - it can not be done unless you get a different, heavier bottomed pan suitable for induction. Exactly what I was saying regarding the replacement costs.

gajjanag··on SF may soon ban natural gas in homes and businesses undergoing major renovations
+1 - there are just so many Asian recipes that can not be done anywhere near as easily on induction stovetops (high heat from direct flame for flatbreads, etc).

Plus a whole bunch of cookware doesn't work with induction (clay pots, non ferromagnetic bases, etc). I do wonder if any of these "environmental" estimates factor in the environmental cost of replacing a bunch of cookware just to satisfy induction requirements.

gajjanag··on Too Many Open Files
There is a vast number of sysctl in xnu that have not really been re-examined in over 15 years. Many tunings date back to the spinning rust drive era (for example). There are plenty of examples like this.

Disclaimer: I worked at Apple and poked xnu a bit.

gajjanag··on Developers, don't despair, big tech and AI hype is off the rails again
The big problem is a bunch of folks actually take these things seriously and use it as an excuse to freeze the junior hiring pipeline.

At the senior levels this is not actually believed by the powers that be, since a bunch of hiring is still happening to compensate for overdone layoffs in spots, etc.

gajjanag··on Career Development: What It Means to Be a Manager, Director, or VP (2015)
> Large corporations believe anyone is replaceable.

This is definitely true. By design, large corporations are structured so that there is no single point of failure.

> Again I am an IC & don’t see/hear any extra work done for retention.

Even in large corporations, extra work definitely happens for retention (I have experienced it myself as an IC). Even though everyone is by design replaceable, the organization has some incentive to work on retention:

a) Bad retention hurts the organization's reputation and future hiring (horror stories spread very fast)

b) Within the team, losing a great teammate hurts morale and output and managers know it will result in a hit on their metrics at least for the next half.

c) Managers may not always be able to backfill, and losing an employee can reduce the size of their "empire" that they are often trying so hard to establish at whatever cost.

gajjanag··on Apple's Software Quality Crisis
Same. The compensation is substantially better at FAANG, but in terms of actual on the ground work being rewarded, almost never the case.

Meta-work (lots of "cross functional" documents, alignment meetings, sync ups with senior tech leads to brown nose, deliberately creating low quality output to justify hiring more people/growing one's "scope") is 90% of it.

Any actual output is largely accidental, coming from the 20% still naive, or idealistic enough to actually care about what they produce.

gajjanag··on An analysis of DeepSeek's R1-Zero and R1
This is much more nuanced now. See Apple "Private Cloud Compute": https://security.apple.com/blog/private-cloud-compute/ ; they run a lot of the larger models on their own servers.

Fundamentally it is more efficient to process a batch of tokens from multiple users/requests than processing them from a single user's request on device.

gajjanag··on "Nvidia is so far ahead that all the 4090s are nerfed to half speed"
Maybe on a particular model/dataset but extremely unlikely in general. Again, like another commenter pointed out: if you truly believe it isn't that hard we would love to hire you at Meta ;)
gajjanag··on Building Meta's GenAI infrastructure
Our group works on some of this stuff at Meta, and we have a pretty good diversity of backgrounds - high performance computing (the bulk), computer systems, compilers, ML engineers, etc. We are hiring.

Feel free to DM me to learn more.

gajjanag··on LLM in a Flash: Efficient LLM Inference with Limited Memory
lmkd (low memory killer daemon) works fairly differently off of a different set of signals and different policy. But yes, conceptually they try to achieve the same goal.

I also do not know if Android combines system libraries into one big file for the savings, something Apple devices do.

gajjanag··on LLM in a Flash: Efficient LLM Inference with Limited Memory
A couple of additional points on how the "low-RAM" works:

1 - https://www.lifewire.com/understanding-compressed-memory-os-... : Apple devices have support for memory compression, see https://opensource.apple.com/source/xnu/xnu-2050.18.24/libke...

2 - Apple devices support something called "jetsam", which basically frees up memory from unused/background apps by killing them in order to keep high priority apps running smoothly: https://developer.apple.com/documentation/xcode/identifying-...

gajjanag··on Size Matters: An Exploration of Virtual Memory on iOS (2022)
Umm, Apple still sells devices with just 1 GB of RAM on them ;)
gajjanag··on Are software developers always forced to do overtime when they miss a deadline?
Same, doing it once out of my own curiosity to see how the corporate machine works.

Not doing it again - seeing first hand how it is due to managerial incompetence more than anything else. The "reward" ratio is just not worth it: if I pull something off; managers will claim its due to their "processes" and "leadership". If I don't pull it off; managers won't promote me.

No win, so... just don't take fake "deadlines" too seriously.

gajjanag··on Google stores billions of lines of code in a single repository (2016) [pdf]
> based on "impact" rather than arbitrary metrics

Umm, from whatever I have seen in big tech "impact" is also fairly arbitrary. It all is based on how cozy one is with one's manager, skip manager, and so on. More accurate is "perception of impact".

Especially as it gets more and more nebulous at higher levels.

gajjanag··on Swift Achieved Dynamic Linking Where Rust Couldn't (2019)
> A page will be loaded in if any part of it is useful. Given that functions will be laid out more or less randomly throughout a shared library, and programs use a randomly scattered subset of the functions, I think its safe to say that you'll get a lot of bytes read in to ram that are never used.

We have order files for this purpose so that functions are not randomly scattered: https://www.emergetools.com/blog/posts/FasterAppStartupOrder... . This technique is widely used by well known apps.

gajjanag··on Introduction to Locality-Sensitive Hashing (2018)
By the way, https://github.com/FALCONN-LIB/FALCONN contains a really good LSH implementation. Also see https://www.mit.edu/~andoni/LSH/ if you want to know more about the research literature.

The fastest way for Euclidean space that I know that works well in practice is via Leech lattice decoding: https://www.semanticscholar.org/paper/Maximum-likelihood-dec... , or https://www.semanticscholar.org/paper/Efficient-bounded-dist... .

It is possible to create an implementation based on the above that decodes 24 dimensional points to the closest Leech lattice vector in < 1 microsecond per point on my AMD Ryzen laptop. Combine with some fast random projections/Johnson Lindenstrauss as described in the article to form the LSH.

This LSH family is unfortunately not present in FALCONN, but the alternatives in FALCONN are pretty good.

Source: thought extensively about LSH for my PhD.

Page 1 of 7Next →