HNHacker News
TopNewBestAskShowJobs

D-Machine

829 karma · joined November 22, 2021

submissionscomments
D-Machine··on Executing programs inside transformers with exponentially faster inference
This is my interpretation as well.

EDIT: Actually, they do make this clear(ish) at the very end of the article, technically. But there is a huge amount of vagueness and IMO outright misleading / deliberately deceptive stuff early on (e.g. about potential differentiability of their approach, even though they admit later they aren't sure if the differentiable approach can actually work for what they are doing). It is hard to tell what they are actually claiming unless you read this autistically / like a lawyer, but that's likely due to a lack of human editing and too much AI assistance.

D-Machine··on Executing programs inside transformers with exponentially faster inference
This is a good link and important (albeit niche) qualification.

It is hard to square with the article's claims about differentiability and otherwise lack of clarity / obscurantism about what they are really doing here (they really are just compiling / encoding a simple computer / VM into a slightly-modified transformer, which, while cool, is really not what they make it sound like at all).

D-Machine··on Executing programs inside transformers with exponentially faster inference
Well, one can never be sure what the real motivation for a lot of DL advances, as most papers are post-hoc obscurantism / hand-waving or even just outright nonsense (see: internal covariate shift explanations for batch norm, which arguably couldn't be more wrong https://arxiv.org/pdf/1805.11604).

When you really get into this stuff, you tend to see the real motivations as either e.g. kernel smoothing (see comments / discussion at https://news.ycombinator.com/item?id=46357675#46359160) or as encoding correlations / feature similarities / multiplicative interactions (see e.g. broad discussion at https://news.ycombinator.com/item?id=46523887). IMO most insights in LLM architectures and layers tends to come from intuitions about projections, manifolds, dimensionality, smoothing/regularization, overparameterization, matrix conditioning, manifold curvature and etc.

There are almost zero useful understandings or insights to be gained from the lookup-table analogy, and most statistical explanations in papers are also post-hoc and require assumptions (convergence rates, infinite layers, etc) that are never shown to clearly hold for actual models that people use. Obviously these AI models work very well for a lot of tasks, but our understanding of why they do is incredibly poor and simplistic, for the most part.

Of course, this is just IMO, and you can see some people in the linked threads do seem to find the lookup table analogies useful. I doubt such people have spent much time building novel architectures, experimenting with different layers, or training such models.

D-Machine··on Executing programs inside transformers with exponentially faster inference
> In our construction, each instruction maps to only a handful of tokens (at most 5).

I don't see how this could work as an LLM given that, but the article is missing a huge amount of other crucial details too.

D-Machine··on Executing programs inside transformers with exponentially faster inference
> i have a pretty good understanding of how transformers work but this did not make sense to me. also i dont understand why this strategy is applicable only to "code tokens"

Yes, there is a monstrous lack of detail here and you should be skeptical about most of the article claims. The language is also IMO non-standard (serious people don't talk about self-attention as lookup tables anymore, that was never a good analogy in the first place) and no good work would just use language to express this, there would also be a simple equation showing the typical scaled dot-product attention formula, and then e.g. some dimension notation/details indicating which matrix (or inserted projection matrix) got some dimension of two somewhere, otherwise, the claims are inscrutable (EDIT: see edit below).

There are also no training details or loss function details, both of which would be necessary (and almost certainly highly novel) to make this kind of thing end-to-end trainable, which is another red flag.

EDIT: The key line seems to be around:

    gate, val = ff_in(x).chunk(2, dim=-1)
and related code, plus the lines "Notice: d_model = 36 with n_heads = 18 gives exactly 2D per head" but, again, this is very unclear and non-standard.
D-Machine··on Executing programs inside transformers with exponentially faster inference
Did you read the post you are responding to? It says:

> What's the benefit? Is it speed? Where are the benchmarks? Is it that you can backprop through this computation? Do you do so?

The correct parsing of this is: "What's the benefit? [...] Is it [the benefit] that you can backprop through this computation? Do you do so?"

There are no details about training nor the (almost-certainly necessarily novel) loss function that would be needed to handle partial / imperfect outputs here, so it is extremely hard to believe any kind of gradient-based training procedure was used to determine / set weight values here.

D-Machine··on Executing programs inside transformers with exponentially faster inference
The post is the perfect example of the kind of writing about AI that dupes people that don't really understand how things like LLMs actually work and are actually trained. Anyone who properly understands these things finds the complete and total lack of detail about training and the loss function (and of course real metrics / benchmarks) to be a monstrous red flag here.

Especially egregious to me is the claim "Because the execution trace is part of the forward pass, the whole process remains differentiable: we can even propagate gradients through the computation itself". This is total weasel-language: e.g. we can propagate any weights through any transformer architecture and all sorts of other much more insane architectural designs, but that is irrelevant if you don't have a continuous and differentiable loss function that can properly weight partially-correct solutions or the likelihood / plausibility of arbitrary model outputs. You also need a clearer source of training data (or way to generate synthetic data).

So for e.g. AlphaFold, we needed to figure out a loss function that continuously approximated the energy configuration of various molecular configurations, and this is what really allowed it to actually do something. Otherwise, you are stuck with slow and expensive reinforcement-based systems.

The other tells are garbage analogies ("Humans cannot fly. Building airplanes does not change that; it only means we built a machine that flies for us"). Such analogies add nothing to understanding, and indeed distract from serious/real understanding. Only dupes and fools think you can gain any meaningful understanding of mathematics and computer science through simplistic linguistic analogies and metaphors without learning the proper actual (visuspatial, logical, etc) models and understanding. Thus, people with real and serious mathematical understanding despise such trite metaphors.

But then, since understanding something like this properly requires serious mathematical understanding, copy like that is a huge tell that the authors / company / platform puts bullshitting and sales above truth and correctness. I.e., yes, a huge yellow flag.

D-Machine··on LLMs work best when the user defines their acceptance criteria first
This article is great. And the blog-article headline is interesting, but wrong. LLM's don't in general write plausible code (as a rule) either.

They just write code that is (semantically) similar to code (clusters) seen in its training data, and which haven't been fenced off by RLHF / RLVR.

This isn't that hard to remember, and is a correct enough simplification of what generative LLMs actually do, without resorting to simplistic or incorrect metaphors.

D-Machine··on A tool that removes censorship from open-weight LLMs
So did I initially until I saw a few more things from others here.
D-Machine··on A tool that removes censorship from open-weight LLMs
Thanks for this link, and mentioning this info some times in this overall thread.

It also seems the influgrifter has a lot of bots (or perhaps cultists) working this thread...

D-Machine··on A tool that removes censorship from open-weight LLMs
Doesn't look legit to me. You are talking about abliteration, which is real. But the OP linked tool is doing novel and very dumb ablation: zeroing out huge components of the network, or zeroing out isolated components in a way that indicates extreme ignorance of the basic math involved.

Compared to abliteration, none of the ablation approaches of this tool make even half a whit of sense if you understand even the most basic aspects of an e.g. Transformer LLM architecture, so my guess is this is BS.

D-Machine··on A tool that removes censorship from open-weight LLMs
You are also not quite correct, IMO. See my comment at https://news.ycombinator.com/item?id=47283197.

What you are talking about is abliteration. What OBLITERATUS seems to be claiming to do is much more dumb, i.e. just zeroing out huge components (e.g. embedding dimension ranges, feed-forward blocks; https://github.com/elder-plinius/OBLITERATUS?tab=readme-ov-f...) of the network as an "Ablation Study" to attempt to determine the semantics of these components.

However, all these methods are marked as "Novel", I.e., maybe just BS made up by the author. IMO I don't see how they can work based on how they are named, they are way too dumb and clunky. But proper abliteration like you mentioned can definitely work.

D-Machine··on A tool that removes censorship from open-weight LLMs
When you look at how monstrously large (and obviously not thought through at all, if you understand even the most minimal basics of the linear algebra and math of a transformer LLM) the components are that are ablated (weights set to zero) in his "Ablation Strategies" section, it is no surprise.

    Strategy            What it does  Use case
    .......................................................
    layer_removal       Zero out      entire transformer layers
    head_pruning        Zero out      individual attention heads
    ffn_ablation        Zero out      feed-forward blocks
    embedding_ablation  Zero out      embedding dimension ranges
https://github.com/elder-plinius/OBLITERATUS?tab=readme-ov-f...
D-Machine··on A tool that removes censorship from open-weight LLMs
I immediately read it as intentional, as a sort of attempt at ironic / nihilistic humour re: LLM-generation, given what the tool claims to do.
D-Machine··on A tool that removes censorship from open-weight LLMs
> "ablation studies", by which it means removing random layers of an already-trained model, to find the source of the refusals(?)

This is not what an ablation study is. An ablation study removes and/or swaps out ("ablates") different components of an architecture (be it a layer or set of layers, all activation functions, backbone, some fixed processing step, or any other component or set of components) and/or in some cases other aspects of training (perhaps a unique / different loss function, perhaps a specialized pre-training or fine-tuning step, etc) in order to attempt to better understand which component(s) of some novel approach is/are actually responsible for any observed improvements. It is a very broad research term of art.

That being said, the "Ablation Strategies" [1] the repo uses, and doing a Ctrl+F for "ablation" in the README does not fill me with confidence that the kind of ablation being done here is really achieving what the author claims. All the "ablation" techniques seem "Novel" in his table [2], i.e. they are unpublished / maybe not publicly or carefully tested, and could easily not work at all.

From later tables, I am not convinced I would want to use these ablations, as they ablate rather huge portions of the models, and so probably do result in massively broken models (as some commenters have noted in this thread elsewhere). EDIT: Also, in other cases [1], they ablate (zero out) architecture components in a way that just seems incredibly braindead if you have even a basic understanding of the linear algebra and dependencies between components of a transformer LLM. There is nothing sound clearly about this, in contrast to e.g. abliteration [3].

[1] hhtps://github.com/elder-plinius/OBLITERATUS?tab=readme-ov-file#ablation-strategies

[2] https://github.com/elder-plinius/OBLITERATUS?tab=readme-ov-f...

EDIT: As another user mentions, "ablation" has a specific additional narrower meaning in some refusal analyses or when looking at making guardrails / changing response vectors and such. It is just a specific kind of ablation, and really should actually be called "abliteration", not "ablation" [3].

[3] https://huggingface.co/blog/mlabonne/abliteration, https://arxiv.org/abs/2512.13655.

D-Machine··on Palantir and other tech companies are stocking offices with tobacco products
> Both of these links go to self-reported data about how addicted people feel themselves to be

This is an incredible and outright lie.

Actually try reading the page I linked, there are plenty of links to scientific studies, scientific reviews, and high-quality resources, as well as lots of careful notes about serious confounds in the usual studies. This includes in exactly the sections I linked.

By all means still be cautious and not careless about using the stuff, that is a perfectly sane position. But I think it is very clear that e.g. patches and gum are highly unlikely to have anything even approaching the risk profile of classic tobacco products.

D-Machine··on Palantir and other tech companies are stocking offices with tobacco products
Plenty of the sources I linked at https://gwern.net/nicotine are scientific and high-quality, that is not just some lazy list or junk compilation.

Nothing wrong with still being cautious about pure nicotine products though, and I definitely would be cautious about vaping.

D-Machine··on Palantir and other tech companies are stocking offices with tobacco products
Correct. My TL:DR summary would be:

- when quantifying for gum and patches, it is really hard to conclude that these pure forms of nicotine are any different from caffeine in harms / addictiveness

- cigarettes, or even pure tobacco smoked is definitely addictive and bad due to other compounds in play

- chewing tobacco products are likely a lot less worse than smoking, but still probably worse than pure nicotine gum or patches

- data is very unclear on vaping

IMO experience suggests vaping can clearly be highly addictive (see: young adults and teens), and I know personally there is even some minimal research on vaping e.g. THC that shows that vaping might be worse than just smoking cannabis, so vaping definitely warrants caution in general [though vaping THC requires much higher temps and different solvents, and involves different terpenes and other compounds].

D-Machine··on Palantir and other tech companies are stocking offices with tobacco products
I think one theory is that nicotine is a vasoconstrictor. Though whether, in its pure form, it is a particularly significant one, i.e. any worse than caffeine, is really not so clear.
D-Machine··on Palantir and other tech companies are stocking offices with tobacco products
An interesting thought, I myself have met at least a couple people that tried to break an addiction by switching to vape products that were essentially just flavour (no nicotine, no THC; THC vapes are common and legal in Canada) and somehow stayed just as addicted to the flavour / oral stimulation. So that sounds at least plausible to me.
D-Machine··on Palantir and other tech companies are stocking offices with tobacco products
Pretty clear from the responses to OP that most people are quite unaware there is almost two decades of decent research on pure nicotine now, and that, outside of vaping (where hard evidence is mostly lacking, due to the novelty), the purer stuff probably really isn't all that addictive, in the grand scheme of things. In many cases it is hard to even say it is much different than caffeine.
D-Machine··on Palantir and other tech companies are stocking offices with tobacco products
Probably not the case for some modern pure nicotine products (gum, patches). Vaping is harder to say due to lack of data and clear cases of addiction in young people, but pure nicotine is definitely a different animal than the classical delivery forms. See my response to GP.
D-Machine··on Palantir and other tech companies are stocking offices with tobacco products
It is really not so clear at all this is the case for pure nicotine products. See https://gwern.net/nicotine#habit-formation, and https://gwern.net/nicotine#dependence for a starter / some brief counter-evidence.
D-Machine··on Palantir and other tech companies are stocking offices with tobacco products
I'm not sure it is actually all that clear that pure nicotine products really are so addictive as people believe. E.g. most studies claiming such addictiveness may simply be because those that get addicted to patches / gum were already addicted to cigarettes (or other classic tobacco product) prior. See e.g. Gwern's notes on the topic.

https://gwern.net/nicotine#habit-formation

https://gwern.net/nicotine#dependence

D-Machine··on Palantir and other tech companies are stocking offices with tobacco products
Seems a good time to link to Gwern's well-researched notes on why nicotine, when consumed in purer forms (e.g. patches, gum), may be pretty useful and not really as harmful nor as addictive as one might think.

https://gwern.net/nicotine

D-Machine··on Government grant-funded research should not be published in for-profit journals
Agree with all this. Once you've filtered / made decisions of quality based on the more substantive criteria, journal reputation can provide useful additional information / context. The case you mentioned is a good example.
D-Machine··on Government grant-funded research should not be published in for-profit journals
I do work in science, I am claiming that pre-publication / journalistic peer review is limiting (and biasing) the amount of post-publication / non-journalistic peer review that can happen, and it is not limiting this in a very reliable or even IMO particularly desirable way.

There is definitely a problem with the over-production of junk science, and we definitely need a way to filter this out somehow. I am just claiming journalistic / pre-publication peer review does not do this effectively or reliably at all anymore (if it ever did).

D-Machine··on Government grant-funded research should not be published in for-profit journals
This is IMO just bad faith sealioning, you can look at the whole replication crisis in psychology and social science (esp. the work of people like Nick Brown and the GRIM test, or Uri Simonsohn), or sites like Retraction Watch, and see clear evidence of everything I am saying. There are endless papers in ML research going into issues with test datasets and data duplication, etc. In plenty of cases all data and code is made open, so it is trivial to check data issues and methods.

Also, review is back and forth, and has rounds: you almost always interrogate the scientists of the paper you are reviewing, this almost like the definition of peer review. I don't think you have any idea of what you are talking about at all.

EDIT: Heck, just hop on over to https://openreview.net/ and take a look at the whole review process for some random paper (e.g. https://openreview.net/forum?id=cp5PvcI6w8_)

D-Machine··on Government grant-funded research should not be published in for-profit journals
I literally said it was posted in this thread, and a quick Ctrl+F of my username on this page would have found you it in a half second: https://news.ycombinator.com/item?id=47249236
D-Machine··on Government grant-funded research should not be published in for-profit journals
> A peer reviewer reads a paper and make comments on it. That's it! They don't check primary data, they don't investigate methods, they don't interrogate scientists, they don't re-run experiments just to double check. They assist a journal's editors in editing--that's it.

Um, what? I have done all these things in reviews, and know other academics that have done these things as well. More confusingly though, if you are saying most reviewers don't do these things (which I agree with), this would only strengthen my point?

I'll let readers decide if it is my comments that exacerbate the problem, or if, perhaps, it is apologism for journalistic peer review that might be causing bigger issues in the present day.

← PreviousPage 7 of 20Next →