HNHacker News
TopNewBestAskShowJobs

D-Machine

831 karma · joined November 22, 2021

submissionscomments
D-Machine··on “Erdos problem #728 was solved more or less autonomously by AI”
Transformer != LLM. See my edited top-level post. Just because Aristotle uses a transformer doesn't mean it is an LLM, just as Vision Transformers and AlphaFold use transformers but are not LLMs.

LLM = Large Language Model. Large refers to both the number of parameters (and in practice, depth) of the model, and also implicitly the amount of data used for training, and "language" means human (i.e. written, spoken) language. A Vision Transformer is not an LLM because it is trained on images, and AlphaFold is not an LLM because it is trained molecular configurations.

Aristotle works heavily with formalized LEAN statements and expressions. While you can certainly argue this is a language of sorts, it is not at all the same "language" as the "language" in LLMs. Calling Aristotle an "LLM" just because it has a transformer is more misleading than truthful, because every other single aspect of it is far more clever and involved.

D-Machine··on “Erdos problem #728 was solved more or less autonomously by AI”
Equating LLMs and transformers is not a meaningless terminology difference at all, Aristotle is so different from the things people call LLMs in terms of training data, loss function, and training that this is a grievous error.
D-Machine··on “Erdos problem #728 was solved more or less autonomously by AI”
You're understanding correctly, this is back and forth between Aristotle and ChatGPT and a (very smart) user.
D-Machine··on “Erdos problem #728 was solved more or less autonomously by AI”
If you think "transformer" = LLM, you don't understand the basic terminology of the field. This is like calling AlphaFold an LLM because it uses a transformer.
D-Machine··on “Erdos problem #728 was solved more or less autonomously by AI”
"Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver"

It is far more than an LLM, and math != "language".

D-Machine··on “Erdos problem #728 was solved more or less autonomously by AI”
"Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver"

It is far more than an LLM, and math != "language".

D-Machine··on “Erdos problem #728 was solved more or less autonomously by AI”
I do / have done research in building deep learning models and custom / novel attention layers, architectures, etc., and AI (ChatGPT) is tremendously helpful in facilitating (semantic) search for papers in areas where you may not quite know the magic key words / terminology for what you are looking for. It is also very good at linking you to ideas / papers that you might not have realized were related.

I also found it can be helpful when exploring your mathematical intuitions on something, e.g. like how a dropout layer might effect learned weights and matrix properties, etc. Sometimes it will find some obscure rigorous math that can be very enlightening or relevant to correcting clumsy intuitions.

D-Machine··on “Erdos problem #728 was solved more or less autonomously by AI”
Did you not read the parts where Aristotle (https://arxiv.org/pdf/2510.01346) was an integral component of all this?
D-Machine··on “Erdos problem #728 was solved more or less autonomously by AI”
It could not be done without Aristotle (https://arxiv.org/pdf/2510.01346), as clearly described in Tao's posts.
D-Machine··on “Erdos problem #728 was solved more or less autonomously by AI”
It could not be done without Aristotle (https://arxiv.org/pdf/2510.01346), did you even read the links?
D-Machine··on Exercise can be nearly as effective as therapy for depression
Dunno why you are being downvoted, probably cope. It is well known by now that antidepressants are only marginally effective on average [1-2]. You're right they should probably only be prescribed for quite severe or treatment-resistant depression. Although the treatment-by-severity effect has been somewhat disputed [3-4], it has rough support [5], and makes sense since it is dubious that we should be giving ineffective medication with serious costs and side-effects to people with moderate depression.

[1] https://ebm.bmj.com/content/27/2/69.abstract

[2] https://ebm.bmj.com/content/25/4/130.abstract

[3] https://link.springer.com/article/10.1186/1744-859X-12-26

[4] https://doi.org/10.1192/bjp.bp.116.187773

[5] https://jamanetwork.com/journals/jama/article-abstract/18515...

D-Machine··on “Erdos problem #728 was solved more or less autonomously by AI”
This is great, there is still so much potential in AI once we move beyond LLMs to specialized approaches like this.

EDIT: Look at all the people below just reacting to the headline and clearly not reading the posts. Aristotle (https://arxiv.org/abs/2510.01346) is key here folks.

EDIT2: It is clear much of the people below don't even understand basic terminology. Something being a transformer doesn't make it an LLM (vision transformers, anyone) and if you aren't training on language (e.g. AlphaFold, or Aristotle on LEAN stuff), it isn't a "language" model.

D-Machine··on Exercise can be nearly as effective as therapy for depression
"Chemical imbalance" theories of depression (e.g. "serotonin hypothesis") have been scientifically discredited for well over a decade now.

EDIT: For a recent overview, see https://www.nature.com/articles/s41380-022-01661-0

D-Machine··on Exercise can be nearly as effective as therapy for depression
You cannot say if this is a substantial change or not, because you need to know by how much the groups actually differ on average, i.e. you need the unstandardized effect size, expressed as a mean difference in the scale sum scores, or as an actual percentage of symptoms reduced, or etc. In general, there are monstrous issues with standardized mean differences, even setting aside the interpretability issues [1-3].

See also my response to GP.

[1] https://journals.plos.org/mentalhealth/article?id=10.1371/jo...

[2] https://bpspsychub.onlinelibrary.wiley.com/doi/abs/10.1348/0...).

[3] https://www.tandfonline.com/doi/abs/10.1080/00031305.2018.15...

D-Machine··on Exercise can be nearly as effective as therapy for depression
Unfortunately, if exercise is only nearly as effective as therapy for depression, it may mean that the benefits of exercise are not actually really clinically observable, if measured properly and not just based on arbitrary statistical significance.

Standardized effect sizes like the ones reported here have no clinical meaning, they are purely statistical. To measure if these kinds of changes matter, you need to determine the Minimal (Clinically) Important Difference [1-2]. I.e. can clinicians (or patients) even notice the observed statistical difference.

In practice, this is a change of about 3-5 points on most 20+ item rating scales, or a relative reduction of 20-30% of the total (sum) score of the scale [1-2]. Unfortunately, anti-depressants are under or just barely reach this threshold [3-4], and so should be widely to be considered ineffective or only borderline effective, on average. Of course this is complicated by the fact that some people get worse on these treatments, and some people experience dramatic improvements, but, still, the point is, depression is extremely hard to treat.

EDIT: There is less data on MCIDs for therapy, but at least one review suggests therapy effects can be in the 10+ point range [5]. But the way the exercise study is presented, with a standardized effect size, we can have no idea if the results matter at all [6].

[1] Button, et al. (2015). Minimal clinically important difference on the Beck Depression Inventory - II according to the patient’s perspective. Psychological Medicine, 45(15), 3269–3279. https://doi.org/10.1017/S0033291715001270 [https://www.cambridge.org/core/journals/psychological-medici...]

[2] Masson, S. C., & Tejani, A. M. (2013). Minimum clinically important differences identified for commonly used depression rating scales. Journal of clinical epidemiology, 66(7), 805-807. [https://www.jclinepi.com/article/S0895-4356(13)00056-5/fullt...]

[3] Hengartner, M. P., & Plöderl, M. (2022). Estimates of the minimal important difference to evaluate the clinical significance of antidepressants in the acute treatment of moderate-to-severe depression. BMJ Evidence-Based Medicine, 27(2), 69-73. https://doi.org/10.1136/bmjebm-2020-111600 [https://ebm.bmj.com/content/27/2/69.abstract]

[4] Jakobsen, J. C., Gluud, C., & Kirsch, I. (2020). Should antidepressants be used for major depressive disorder?. BMJ evidence-based medicine, 25(4), 130-130. https://doi.org/10.1136/bmjebm-2019-111238 [https://ebm.bmj.com/content/25/4/130.abstract]

[5] Cuijpers, P., Karyotaki, E., Weitz, E., Andersson, G., Hollon, S. D., & van Straten, A. (2014). The effects of psychotherapies for major depression in adults on remission, recovery and improvement: a meta-analysis. Journal of affective disorders, 159, 118–126. https://doi.org/10.1016/j.jad.2014.02.026 [https://pubmed.ncbi.nlm.nih.gov/24679399/]

[6] Pogrow, S. (2019). How Effect Size (Practical Significance) Misleads Clinical Practice: The Case for Switching to Practical Benefit to Assess Applied Research Findings. The American Statistician, 73(sup1), 223–234. https://doi.org/10.1080/00031305.2018.1549101

D-Machine··on Exercise can be nearly as effective as therapy for depression
It means nothing, standardized effect sizes have no clinical meaning here, they are purely statistical. To measure if these kinds of changes matter, you need to determine the Minimal (Clinically) Important Difference [1-2]. I.e. can clinicians (or patients) even notice the observed statistical difference.

In practice, this is a change of about 3-5 points on most 20+ item rating scales, or a relative reduction of 20-30% of the total (sum) score of the scale [1-2]. Unfortunately, anti-depressants are under or just barely reach this threshold [3-4], and so should be widely to be considered ineffective or only borderline effective, on average. Of course this is complicated by the fact that some people get worse on these treatments, and some people experience dramatic improvements, but, still, the point is, depression is extremely hard to treat.

Unfortunately, this also means that if exercise is only nearly as effective as therapy for depression, it may mean that the benefits of exercise are not actually really clinically observable, if measured properly and not just based on arbitrary statistical significance.

EDIT: There is less data on MCIDs for therapy, but at least one review suggests therapy effects can be in the 10+ point range [5]. But the way the exercise study is presented, with standardized effect sizes, we have no idea if the results matter at all [6].

[1] Button, et al. (2015). Minimal clinically important difference on the Beck Depression Inventory - II according to the patient’s perspective. Psychological Medicine, 45(15), 3269–3279. https://doi.org/10.1017/S0033291715001270 [https://www.cambridge.org/core/journals/psychological-medici...]

[2] Masson, S. C., & Tejani, A. M. (2013). Minimum clinically important differences identified for commonly used depression rating scales. Journal of clinical epidemiology, 66(7), 805-807. [https://www.jclinepi.com/article/S0895-4356(13)00056-5/fullt...]

[3] Hengartner, M. P., & Plöderl, M. (2022). Estimates of the minimal important difference to evaluate the clinical significance of antidepressants in the acute treatment of moderate-to-severe depression. BMJ Evidence-Based Medicine, 27(2), 69-73. https://doi.org/10.1136/bmjebm-2020-111600 [https://ebm.bmj.com/content/27/2/69.abstract]

[4] Jakobsen, J. C., Gluud, C., & Kirsch, I. (2020). Should antidepressants be used for major depressive disorder?. BMJ evidence-based medicine, 25(4), 130-130. https://doi.org/10.1136/bmjebm-2019-111238 [https://ebm.bmj.com/content/25/4/130.abstract]

[5] Cuijpers, P., Karyotaki, E., Weitz, E., Andersson, G., Hollon, S. D., & van Straten, A. (2014). The effects of psychotherapies for major depression in adults on remission, recovery and improvement: a meta-analysis. Journal of affective disorders, 159, 118–126. https://doi.org/10.1016/j.jad.2014.02.026 [https://pubmed.ncbi.nlm.nih.gov/24679399/]

[6] Pogrow, S. (2019). How Effect Size (Practical Significance) Misleads Clinical Practice: The Case for Switching to Practical Benefit to Assess Applied Research Findings. The American Statistician, 73(sup1), 223–234. https://doi.org/10.1080/00031305.2018.1549101

D-Machine··on The Q, K, V Matrices
1. I do not think it is orthogonal, but, regardless, there is plenty of research trying to get explainability out of all aspects of scaled dot-product attention layers (weights, QKV projections, activations, other aspects), and trying to explain deep models generally via sort of bottom-up mechanistic approaches. I think it can be clearly argued this does not give us much and is probably a waste of time (see e.g. https://ai-frontiers.org/articles/the-misguided-quest-for-me...). I think this is especially clear when you have evidence (in research, at least) that other mechanisms and layers can produce highly similar results.

2. I didn't say the transformations can be reversed, I said if you interpret anything as an importance (e.g. a magnitude), that can be inflated / reversed by whatever weights are learned by later layers. Negative values and/or weights make this even more annoying / complicated.

3. Not sure how this is relevant, but, yes, any reasons for caring about QKV and scaled dot-product attention specifics are mostly related to performance and/or current popular leading models. But there is nothing fundamentally important about scaled dot-product attention, it most likely just happens to be something that was settled on prematurely because it works quite well and is easy to parallelize. Or, if you like the kernel smoothing explanation also mentioned in this thread, scaled dot-product self-attention implements something very similar to a particularly simple and nice form of kernel smoothing.

4. Yup, removing ops from scaled dot-product attention blocks is going to dramatically reduce expressivity, because there really aren't much ops there to remove. But there is enough work on low-rank attention, linear attentions, and sparse attentions, that show you can remove a lot of expressivity and still do quite well. And, of course, the huge amount of helpful other types of attention I linked before give gains in some cases too. You should be skeptical about any really simple or clear story about what is going on here. In particular, there is no clear reason why a small hypernetwork couldn't be used to approximate something more general than scaled dot-product attention, except that, obviously this is going to be more expensive, and in practice you can probably just get the same approximate flexibility by stacking simpler attention layers.

5. I still find that doesn't give me any clear mathematical meaning.

I suspect our learning goals are at odds. If you want to focus solely on the very specific kind of attention used in the popular transformer models today, perhaps because you are interested in optimizations or distillation or something, then by all means try to come up with special intuitions about Q, K, and V, if you think that will help here. But those intuitions will likely not translate well to future and existing modifications and improvements to attention layers, in transformers or otherwise. You will be better served learning about attention broadly and developing intuitions based on that.

Others have mentioned the kernel smoothing interpretation, and I think multiplicative interactions are the clearer deeper generalization of what is really important and valuable here. Also, the useful intuitions in DL have been less about e.g. "feature importances" and "sensitivity" and such, but tend to come more from linear algebra and calculus, and tend to involve things like matrix conditioning and regularization / smoothing and Lipschitz constants and the like. In particular, the softmax in self-attention is probably not doing what people typically say it does (https://arxiv.org/html/2410.18613v1), and the real point is that all these attention layers are trained in an end-to-end fashion where all layers are interdependent on each other to varying complicated degrees. Focusing on very specific interpretations ("Q is this, K is that"), especially where these interpretations are sort of vaguely metaphorical, like yours, is not likely to result in much deep understanding, in my opinion.

D-Machine··on The Q, K, V Matrices
In literary and casual contexts, absolutely (though we'd probably say "he/she" instead of "it" here). As I said, "it" referring to the mat is the most natural and obvious reading, but other ones are perfectly logical and sound, if less likely/common.

Although the sentence is itself a bit awkward and strange on its own, and really needs context. In fact, this is because the sentence is generated as a short example to make a point about attention and tokens, and is not really something someone would utter naturally in isolation.

I mostly just wanted to playfully comment that original GP / top-level comment had a valid point about the ambiguity!

D-Machine··on Sugar industry influenced researchers and blamed fat for CVD (2016)
Nutrition science is not science in almost any of the ways a real science needs to be, and there is almost zero "real, good science" to be found in it. The reasons this statement is true (as well as the precise qualifications of the exceptions to this) are well laid out by tsimionescu in response to your post.

The measurement, control, confounds, and even basic concepts are atrocious here, this is possibly the only field as bad as or even worse than e.g. social psychology. And this is all ignoring the massive economic interests involved.

It is in fact only science illiteracy that would lead one to think nutrition science is a serious science. At the most absolute charitable, it is a protoscience like alchemy (which did have some replicable findings that eventually led to real chemistry, but which was still mostly nonsense at core).

D-Machine··on Sugar industry influenced researchers and blamed fat for CVD (2016)
This. The amount of faith in nutrition "science" indicates severe science illiteracy in the public.

In general there are way too many confounds, and measurement is far too poor and unreliable (self-report that is wrong in quality and quantity; you can't track enough people for the amount of time where supposed effects would manifest), there is almost zero control over what people eat (diets and available foods even considerably over a decade for whole countries, never mind within individuals), and much of the things being measured lack even face/content validity in the first place (e.g. "fat" is not a valid category, and even "saturated vs. unsaturated" is a matter of degree, and each again with different kinds in each category).

We are missing so much of the basics of what are required for a real science here I think it is far more reasonable to view almost all long-term nutritional claims as pseudoscience, unless the effect is clear and massive (e.g. consumption of large amounts of alcohol, or extremely unique / restrictive diets that have strong effects, or the rare results of natural experiments / famines), or so extremely general that it catches a sort of primary factor (too much calories is generally harmful, regardless of the source of those calories).

Maybe it'll become actual science one day, but that won't be for decades.

D-Machine··on Sugar industry influenced researchers and blamed fat for CVD (2016)
Fair, the term may have been well-defined and measured in the original study, or in some specific circles. I was definitely thinking of the meaningless general thing "Mediterranean diet" has metastasized into today.

I also think it is better, rhetorically, to not draw support for the badness of saturated fats / differences of different fats by referencing the Mediterranean diet, since this rather looks like drawing upon narrow / weak science to support something that is in fact much more broadly supported by a larger variety of more careful work.

But yes, it is very important that people recognize there are huge differences here!

D-Machine··on The Q, K, V Matrices
Ah, but cats won't just comfortably sit on a mat if they feel there is danger. They will only sit on a mat if they feel comfortable! Absent larger context, the sentence is in fact ambiguous (though I agree your reading is the most natural and obvious one).
D-Machine··on Sugar industry influenced researchers and blamed fat for CVD (2016)
Epidemiology should generally be disregarded when it comes to nutrition.

There are exceptions when there are rare natural experiments (e.g. I forget the country, but the European one where some issue caused all flour for the country to be only whole-wheat, which led to clear nutrient deficiencies due to the phytic acid there) but in general there are way too many confounds, and measurement is far too poor and unreliable (self-report that is not just quantitatively but qualitatively wrong, and you can't track enough people nearly long enough), there is virtually no control whatsoever (diets and available foods shift considerably over just decades), and much of the things being measured lack even face/content validity in the first place (e.g. "fat" is not a valid taxon, and even "saturated vs. unsaturated" is a matter of degree).

We are missing so much of the basics of what are required for a real science here I think it is far more reasonable to view almost all long-term nutritional claims as pseudoscience, unless the effect is clear and massive (e.g. consumption of large amounts of alcohol, or extremely unique / restrictive diets that have strong effects), or so extremely general that it catches a sort of primary factor (too much calories is generally harmful, regardless of the source of those calories).

But even setting that aside, you can't define or study "Mediterranean diet" rigorously even in RCTs, so I don't see how you can think you are going to get much of anything here from epidemiological work that is going to lead to anything practically actionable.

D-Machine··on The Q, K, V Matrices
Read through my comments and those of others in this thread, the way you are thinking here is metaphorical and so disconnected from the actual math as to be unhelpful. It is not that case that you can gain a meaningful understanding of deep networks by metaphor. You actually need to learn some very basic linear algebra.

Heck, attention layers never even see tokens. Even the first self-attention layer sees positional embeddings, but all subsequent attention layers are just seeing complicated embeddings that are a mish-mash of the previous layers' embeddings.

D-Machine··on The Q, K, V Matrices
Read the third link / review paper, it is not at all the case that all attention is based on QKV projections.

Your terms "sensitivity", "visibility", and "important" are too vague and lack any clear mathematical meaning, so IMO add nothing to any understanding. "Important" also seems factually wrong, given these layers are stacked, so later weights and operations can in fact inflate / reverse things. Deriving e.g. feature importances from self-attention layers remains a highly disputed area (e.g. [1] vs [2], for just the tip of the iceberg).

You are also assuming that the importance of attention is the highly-specific QKV structure and projection, but there is very little reason to believe that based on the third review link I shared. Or, if you'd like another example of why not to focus so much on scaled dot-product attention, see that it is just a subset of a broader category of multiplicative interactions (https://openreview.net/pdf?id=rylnK6VtDH).

[1] Attention is not Explanation - https://arxiv.org/abs/1902.10186

[2] Attention is not not Explanation - https://arxiv.org/abs/1908.04626

D-Machine··on Sugar industry influenced researchers and blamed fat for CVD (2016)
Mediterranean diet is nonsense. Ill-defined, doesn't have clear evidence of a relation to CVD in hard studies. Bad that people still believe this.

https://pmc.ncbi.nlm.nih.gov/articles/PMC6414510/

D-Machine··on Sugar industry influenced researchers and blamed fat for CVD (2016)
Cochrane systematic reviews should make you seriously question whether the Mediterranean diet really is much good at all - hard data is inconclusive and low quality [1].

In general we really even barely have enough nutritional knowledge to say if the term 'good fats' even makes much scientific sense, but broad and vague things like "Mediterranean diet" are just total nonsense, from the standpoint of serious nutrition science.

[1] https://pmc.ncbi.nlm.nih.gov/articles/PMC6414510/

D-Machine··on Sugar industry influenced researchers and blamed fat for CVD (2016)
You can definitely get to 20% without much trouble, maybe even 30-50%, if you do some freezing tricks. Though what such percentages mean is highly subjective.

I am thinking if an ideal butter croissant has some flaky fluffiness (perhaps if we define it as "trapped volume" between flakes), and we define this ideal flakiness to be 100%, then you can extremely easily get to 20% with just olive oil. Frankly I think you might even get close to 50% (defined in this way) provided you also start with a trustworthy recipe by mass and that aims for proper hydration (e.g. https://www.seriouseats.com/croissants-recipe-11863500) and work quickly with lots of chilling.

Just, subjectively, you might realize that 20-50%, defined this way, isn't much like a proper French croissant, and is more like a cheap doughy supermarket chain croissant—which I do still frankly enjoy sometimes anyway!

D-Machine··on The Q, K, V Matrices
Right, I think there are plenty of other approaches that surely scale just as easily or better. It's like you said, the (early) dominance of text data just artificially narrowed the approaches tried.
D-Machine··on The Q, K, V Matrices
I think this mostly comes down to (multi-headed) scaled dot-product attention just being very easy to parallelize on GPUs. You can then make up for the (relative) lack of expressivity / flexibility by just stacking layers.
← PreviousPage 14 of 20Next →