HNHacker News
TopNewBestAskShowJobs

D-Machine

831 karma · joined November 22, 2021

submissionscomments
D-Machine··on The Q, K, V Matrices
The main takeaway is that "attention" is a much broader concept generally, so worrying too much about the "scaled dot-product attention" of transformers deeply limits your understanding of what kinds of things really matter in general.

A paper I found particularly useful on this was generalizing even farther to note the importance of multiplicative interactions more generally in deep learning (https://openreview.net/pdf?id=rylnK6VtDH).

EDIT: Also, this paper I was looking for dramatically generalizes the notion of attention in a way I found to be quite helpful: https://arxiv.org/pdf/2111.07624

D-Machine··on The Q, K, V Matrices
Don't get caught up in interpreting QKV, it is a waste of time, since completely different attention formulations (e.g. merged attention [1]) still give you the similarities / multiplicative interactions, but may even work better [2]. EDIT: Oh and attention is much more broad than scaled dot-product attention [3].

[1] https://www.emergentmind.com/topics/merged-attention

[2] https://blog.google/innovation-and-ai/technology/developers-...

[3] https://arxiv.org/abs/2111.07624

D-Machine··on The Q, K, V Matrices
Yes, this needs to be linked more, you are doing a great service.
D-Machine··on The Q, K, V Matrices
> I've never really got it and I just switched to thinking of QKV as a way to construct a fairly general series of linear algebra transformations on the input of a sequence of token embedding vectors x that is quadratic in x and ensures that every token can relate to every other token in the NxN attention matrix.

That's because what you say here is the correct understanding. The lookup thing is nonsense.

The terms "Query" and "Value" are largely arbitrary and meaningless in practice, look at how to implement this in PyTorch and you'll see these are just weight matrices that implement a projection of sorts, and self-attention is always just self_attention(x, x, x) or self_attention(x, x, y) in some cases (e.g. cross-attention), where x and y are are outputs from previous layers.

Plus with different forms of attention, e.g. merged attention, and the research into why / how attention mechanisms might actually be working, the whole "they are motivated by key-value stores" thing starts to look really bogus. Really it is that the attention layer allows for modeling correlations/similarities and/or multiplicative interactions among a dimension-reduced representation. EDIT: Or, as you say, it can be regarded as kernel smoothing.

D-Machine··on The Q, K, V Matrices
Very much this, cross attention and the x, y notation makes the similarity / covariance matrix far more clear and intuitive.

Also forget the terms "query", "key" and "value", or vague analogies to key-value stores, that is IMO a largely false analogy, and certainly not a helpful way to understand what is happening.

D-Machine··on Opus 4.5 is not the normal AI agent experience that I have had thus far
I have to agree, this video is hardly what most people would mean by programming. I am sure there are better videos than this?
D-Machine··on Sugar industry influenced researchers and blamed fat for CVD (2016)
Butter is NOT favored because most people had it in their youth, but because of its extremely distinct flavour.

"Unctuous" is certainly not specific enough, the reason butter (and ghee) is so delicious is its butteriness, i.e. it has a highly distinct taste. All properly rendered animal fats have highly distinct tastes and serve different purposes. Schmaltz tastes slightly of chicken, duck fat of duck, lard of pork, and tallow of beef.

But butter does NOT distinctly taste of beef, rather, it is reminiscent of slightly-aged milk (or, in the case of ghee, it may even strongly smell like certain kinds of aged cheese). There is, also, in butter, significant absorbed water content, and, to my palate, even a very subtle acidity that is not quite present in other rendered animal fats that give it a sort of brightness that make it work in things like butter-creams and other delicate or mild flavours (e.g. popcorn).

It is IMO this specifically "non-meaty" unctuousness that is the real draw of butter. Not some childhood nostalgia.

D-Machine··on Sugar industry influenced researchers and blamed fat for CVD (2016)
Yes, lamination can be done with almost any fat, but the more you laminate (more layers / folds), the more that liquid fats sort of absorb into the dough, and stop having the desired separating effect. So while oil layering works well for e.g. paratha-style roti, scallion pancakes, and things that only really get one or two "layers" or "folds", oil is just fine. But when you get to something like a croissant, or even just a rough puff pastry (e.g. https://www.seriouseats.com/old-fashioned-flaky-pie-dough-re...), liquid fats are usually a complete non-starter.

You might be able to achieve something if you can somehow freeze your olive oil and chill your dough, and work very quickly during lamination, but you should, even with a lot of work and tweaking, still expect to get a noticeably inferior product for something like croissants.

Depending on how picky you are/not, you might still be personally happy with the texture and taste, but don't expect to get even remotely close to an actual good butter croissant, by more objective standards. Here in Canada we had a minor problem with the butter texture due to what we feed our cows here ("buttergate"), and this was preventing professional bakers from achieving quality croissants with just the Canadian butter. This should make you highly skeptical that you can get anything good with something as different as olive oil.

Still, I do love the idea of an olive oil croissant, it would be delicious.

D-Machine··on Opus 4.5 is not the normal AI agent experience that I have had thus far
Also obviously brains are both!
D-Machine··on Opus 4.5 is not the normal AI agent experience that I have had thus far
You won't find any trustworthy papers on the topic because GP is simply wrong here.

That models can be distilled has no bearing whatsoever on whether a model has learned actual knowledge or understanding ("logic"). Models have always learned sparse/approximately-sparse and/or redundant weights, but they are still all doing manifold-fitting.

The resulting embeddings from such fitting reflect semantics and semantic patterns. For LLMs trained on the internet, the semantic patterns learned are linguistic, which are not just strictly logical, but also reflect emotional, connotational, conventional, and frequent patterns, all of which can be illogical or just wrong. While linguistic semantic patterns are correlated with logical patterns in some cases, this is simply not true in general.

D-Machine··on Opus 4.5 is not the normal AI agent experience that I have had thus far
Thank you for linking this very useful and much more realistic / grounded stat.
D-Machine··on Travel Is Not Education
Nothing about my post said nor implied the two things were at odds or that there aren't people that do both. In fact, based on everything you are saying, I can't really find anything to disagree with at all.

> So there’s travel and there’s travel.

Indeed. If travel = tourism, then I agree most travel (as tourism) is superficial gives mostly trivial knowledge about a culture. If travel is "living / working abroad" or "an exchange", than, obviously it is not so trivial. And indeed, even a week as a tourist can be rich if you've read deeply on some specific aspect of the country, and that is the focus of your tourism.

I would still guess that over 90% of travel (at least among younger generations) is just shallow tourism, and people most vocal about the benefits of travel are generally just tourists pretending their shallow tourism is something more. This is the sentiment I think animates this kind of blog post / article.

EDIT: And also there is nothing wrong with liking fundamentally superficial and/or simple things. I enjoy trashy fast food and SPAM from time to time despite also happily spending many days and hours carefully preparing gourmet meals. But I don't ever pretend that enjoying SPAM is some elevated fine taste. Those who enjoy shallow tourism just have an annoying tendency to try to pretend that their "travel" somehow makes them better and/or sophisticated in some way, but, it simply does not, in the vast majority of cases.

D-Machine··on Opus 4.5 is not the normal AI agent experience that I have had thus far
> 2. LLMs are far, far more efficient than humans in terms of resource consumption for a given task: https://www.nature.com/articles/s41598-024-76682-6 and https://cacm.acm.org/blogcacm/the-energy-footprint-of-humans...

I want to push back on this argument, as it seems suspect given that none of these tools are creating profit, and so require funds / resources that are essentially coming from the combined efforts of much of the economy. I.e. the energy externalities here are monstrous and never factored into these things, even though these models could never have gotten off the ground if not for the massive energy expenditures that were (and continue to be) needed to sustain the funding for these things.

To simplify, LLMs haven't clearly created the value they have promised, but have eaten up massive amounts of capital / value produced by everyone else. But producing that capital had energy costs too. Whether or not all this AI stuff ends up being more energy efficient than people needs to be measured on whether AI actually delivers on its promises and recoups the investments.

EDIT: I.e. it is wildly unclear at this point that if we all pivot to AI that, economy-wide, we will produce value at a lower energy cost, and, even if we grant that this will eventually happen, it is not clear how long that will take. And sure, humans have these costs too, but humans have a sort of guaranteed potential future value, whereas the value of AI is speculative. So comparing energy costs of the two at this frozen moment in time just doesn't quite feel right to me.

D-Machine··on Travel Is Not Education
All true, but most "travel" is staying a week or so as a tourist in some location, and it is true that what is learned from this kind of travel is generally trivial and superficial (and thus often wrong). You probably can learn more deep truths about a country and culture from reading on the internet, unless you are really making an effort to properly integrate in some way, likely for a minimum of many months. But, then, this is usually not what is meant by "travel".
D-Machine··on DHH: AI models are now good enough
Great example of something that actually has some substance beyond meaningless anecdotes.
D-Machine··on LeCun calls Alex Wang inexperienced, predicts more Meta AI employee departure
Agreed.
D-Machine··on DHH: AI models are now good enough
What can be stated without evidence can be dismissed without evidence. It is IMO pretty clear to me there is no substance to this post, without knowing anything about the author.

In general most such claims today are without substance, as they are made without any real metrics, and the metrics we actually need we just don't have. I.e. we need to quantify the technical debt of LLM code, how often it has errors relative to human-written code, and how critical / costly those errors are in each case relative to the cost of developer wages, and also need to be clear if the LLM usage is just boilerplate / webshit vs. on legacy codebases involving non-trivial logic and/or context, and whether e.g. the velocity / usefulness of the LLM-generated code decreases as the codebase grows, and etc.

Otherwise, anyone can make vague claims that might even be in earnest, only to have e.g. studies show that in fact the productivity is reduced, despite the developer "feeling" faster. Vague claims are useless at this point without concrete measurements and numbers.

D-Machine··on Boredom Is Good, Actually
Dead link: Failed Dependency (Error 424)!
D-Machine··on LeCun calls Alex Wang inexperienced, predicts more Meta AI employee departure
This too should be questioned, at least a couple studies at this point suggesting many feel like they are going faster with AI when, by some metrics, they are going slower (e.g. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-o...), and then there are e.g. admissions from major CEOs publicly admitting e.g. Copilot doesn't "really work" (https://ppc.land/microsoft-ceo-admits-copilot-integrations-d...).

And, again, this is ignoring all the technical debt of produced code that is poorly understood, weakly-reviewed, and of questionable quality overall.

I still think this all has serious potential for net benefit, and does now in certain cases. But we need to be clearer about spelling out where that is (webshit, boilerplate, language-to-language translation, etc) and where it maybe isn't (research code, legacy code, large codebases, niche/expert domains).

D-Machine··on LeCun calls Alex Wang inexperienced, predicts more Meta AI employee departure
Or, people who've actually trained and used models in domains where "stuff on the internet" is of no relevance to what you are actually doing realize the profound limitations to what these LLMs actually do. They are amazing, don't get me wrong, but not so amazing in many specific contexts.
D-Machine··on LeCun calls Alex Wang inexperienced, predicts more Meta AI employee departure
Yup, most progress is also confined to SWE's doing webshit / writing boilerplate code too. Anything specialized, LLMs are rarely useful, and this is all ignoring the future technical debt of debugging LLM code.

I am hopeful about LLMs for SWE, but the progress is currently contextual.

D-Machine··on LeCun calls Alex Wang inexperienced, predicts more Meta AI employee departure
It isn't wrong, just think about how weights are updated via (mini-)batches, and how tokenization works, and you will understand that LLM's can't ignore poisoning / outliers like humans do. This would be a classic recent example (https://arxiv.org/abs/2510.07192): IMO because the standard (non-robust) loss functions allow for anchor points .
D-Machine··on LeCun calls Alex Wang inexperienced, predicts more Meta AI employee departure
Your best bet would be to look deeply into performance on ARC-AGI fully-private test set performances (e.g. https://arcprize.org/blog/arc-prize-2025-results-analysis), and think carefully about the discrepancies here, or, just to broadly read any academic research on classic benchmarks and note the plateaus on classic datasets.

It is very clear when you look at academic papers actually targeting problems specific to reasoning / intelligence (e.g. rotation invariance in images, adversarial robustness) that all the big companies are doing is just fitting more data / spending more resources on human raters and other things to boost performance on (open) metrics, but that clear actual gains in genuine intelligence are being made only by milking what we know very well to be a limited approach. I.e. there are trivially-basic problems that cannot be solved by curve-fitting models, which makes it clear most current advances are indeed coming from curve(manifold) fitting. It just isn't clear how far we can exploit these current approaches and in what domains this kind of exploitation is more than good enough.

EDIT: Are people unaware Google Scholar is a thing? It is trivial to find modern AI papers that can be read without requiring access to a research institution. And e.g. HuggingFace collects trending papers (https://huggingface.co/papers/trending), and etc.

D-Machine··on LeCun calls Alex Wang inexperienced, predicts more Meta AI employee departure
I should clarify that LLMs trained on the internet are necessarily a dead end, theoretically, because the internet both (1) lacks specialist knowledge and knowledge that cannot be encoded in text / language, and (2) is polluted with not just false, but irrelevant knowledge for general tasks. LLMs (or rather, transformers and deep models tuned by gradient descent) trained on synthetic data or more curated / highly-specific data where there are actual costs / losses we can properly model (e.g. AlphaFold) could still have tremendous potential. But "LLM" in the usual, everyday sense in which people use this label, are very limited.

A good example would be trying to make an LLM trained on the entire internet do math proofs. Almost everything in its dataset tells it that the word "orthogonal" means "unrelated to", because this is how it is used colloquially. Only in a tiny amount of math forums / resources it digested does this actually mean something about the dot product, so clearly an LLM that does math well only does so by ignoring the majority of the space it is trained on. Similar considerations apply for attempting to use e.g. vision-language models trained on "pop" images to facilitate the analysis of, say, MRI scans, or LIDAR data. That we can make some progress in these domains tells us there is some substantial overlap in the semantics, but it is obvious there are limits to this.

There is no reason to believe these (often: irrelevant, incorrect) semantics learned from the entire web are going to be helpful for the LLM to produce deeply useful math / MRI analysis / LIDAR interpretation. Broadly, not all semantics useful in one domain are useful in another, and, even more clearly, linguistic semantics clearly have limited relevance to much of what we consider intelligence (which includes visual, auditory, proprioceptive/kinaesthetic, and, arguably, mathematical abstractions). But, it could well be that curve-fitting huge amounts of data from the relevant semantic space (e.g. feeding transformers enough Lean / MRI / LIDAR data) is in fact all we need, so that e.g. transformers are "good enough" for achieving most basic AI aims. It just is clearly the case that the internet can't provide all that data for all / most domains.

EDIT: Also Anthropic's writeups are basically fraud if you actually understand the math, there is no "thinking ahead" or "planning in advance" in any sense, literally just if you head down certain paths due to pre-training, yes, of course, you can "already see" weight activations of future tokens: this is just what curve-fitting in N-D looks like, there is no where else for the model to go. Actual thinking ahead means things like backtracking / backspace tokens, i.e. actually retracing your path, which current LLMs simply cannot do.

D-Machine··on LeCun calls Alex Wang inexperienced, predicts more Meta AI employee departure
If you are paying attention to actual research, guarded benchmarks, and understand how benchmarks are being gamed, I would say there is plenty of evidence we are approaching a clear plateau / the march-of-nines thesis of Karpathy is basically correct long-term. Short-term it remains to be seen how much more we can do with the current tech.
D-Machine··on LeCun calls Alex Wang inexperienced, predicts more Meta AI employee departure
> Not sure if anyone who works in the foundational model space and who doesn't directly depend on LLMs 'making it' for VC money would claim differently

This is the problem. The vast majority of people over-hyping LLMs don't even have the most basic understanding of how simple LLMs are at core (manifold-fitting the semantic space of the internet), and so can't understand why they are necessarily dead ends, theoretically. This really isn't debatable for anyone with a full understanding of the training + basic dynamics of what these models do.

But, practically, it remains to be seen where the dead end with LLMs lies. I think we are clearly approaching plateaus in both academic research and in practice (people forget or are unaware how much benchmarks are being gamed as well), but, even small practical gains remain game-changers in this space, and much of the progress / tradeoffs we actually care about can't be measured accurately yet (e.g. rapid development vs. "technical debt" from fast but not-understood / weakly-reviewed LLM code).

LLMs are IMO undebatably a theoretical dead end, and for that reason, a practical dead end too. But we haven't hit that practical dead end yet.

D-Machine··on LeCun calls Alex Wang inexperienced, predicts more Meta AI employee departure
He's right that LLMs are a dead end, but yeah, those quotes were cringe as hell. Hubris.
D-Machine··on Go Gray, Not Cray: Why You Should Grayscale Your Phone
They said they noticed themselves looking at their phone far less. Most likely this is what saved the power, just a typical spurious correlation and a bad theory for the reason.
D-Machine··on Go Gray, Not Cray: Why You Should Grayscale Your Phone
They said they found themselves using their phone far less due to the grayscale, which would be the real thing extending battery life here. Or at least, this was what I assumed on reading.
D-Machine··on How we lost communication to entertainment
The more typical / appropriate word here would be "over-dramatic" or "melodramatic", not "dramatic".
← PreviousPage 15 of 20Next →