HNHacker News
TopNewBestAskShowJobs

D-Machine

829 karma · joined November 22, 2021

submissionscomments
D-Machine··on Ontario auditors find doctors' AI note takers routinely blow basic facts
GP is obviously wrong, and probably doesn't know about calibration and/or that it isn't even clear how to calibrate frontier models in the manner we need, given how complex and expensive the training is, and how tricky calibration becomes in e.g. mixture-of-experts and chain of thought approaches.
D-Machine··on Ontario auditors find doctors' AI note takers routinely blow basic facts
Common misconception. As far we know, LLMs are not calibrated, i.e. their output "probabilities" are not in fact necessarily correlated with the actual error rates, so you can't use e.g. the softmax values to estimate confidence. It is why it is more accurate to talk about e.g. the model "logits", "softmax values", "simplex mapping", "pseudo-probabilities", or even more agnostically, just "output scores", unless you actually have strong evidence of calibration.

To get calibrated probabilities, you actually need to use calibration techniques, and it is extremely unclear if any frontier models are doing this (or even how calibration can be done effectively in fancy chain-of-thought + MoE models, and/or how to do this in RLVR and RLHF based training regimes). I suppose if you get into things like conformal prediction, you could ensure some calibration, but this is likely too computationally expensive and/or has other undesirable side-effects.

EDIT: Oh and also there are anomaly detection approaches, which attempt to identify when we are in outlier space based on various (e.g. distance) metrics based on the embeddings, but even getting actual probabilities here is tricky. This is why it is so hard to get models to say they "don't know" with any kind of statistical certainty, because that information isn't generally actually "there" in the model, in any clean sense.

D-Machine··on Is my blue your blue? (2024)
Except you can reject the very (stupid) question / framing, in which case, the response is to either close the tab, or respond in a particular response style, neither of which makes the data more informative. This kind of clumsy stuff is just dumb with what we know now, edutainment distraction for the HN crowd.
D-Machine··on Is my blue your blue? (2024)
Yes, very annoying, we know from extensive work in psychometrics that single-item, binary / forced-choice items produce junk responses that are heavily contaminated with response styles (answer in most socially-desirable way, select closest response to mouse/finger, select same response as last time, select random response, etc). Just give people an out ("Diagree with the question / premises", "Prefer not to answer", "Unsure / Can't decide", etc) and make sure you have e.g. a 5-7 point Likert-type scale for multiple items, or up to an 11-point scale for single items.

This kind of site / demo does none of the above, and so can't even be trusted for directional effects (the direction of response may simple be due to the type of people responding, etc).

D-Machine··on Is my blue your blue? (2024)
Forced binary choices on single-item, self-report questions produce scientific junk, absolutely. This kind of design / approach encourages not only magnitude errors, but also sign errors (you can't even trust the direction of the observed effect).

IMO, growing up, unless you lived under a rock, it seems obvious to me that you will have experienced different people pointing at the same colour and uttering very different colour labels (pink vs. red, blue vs. green, black vs. deep blue/purple, etc) from the labels you might have applied yourself. Differing/shared colour perception isn't exactly a rare kind of topic (almost is like the canonical stoner topic, also common online), so I'd be a bit surprised if this demo is actually introducing anyone to this concept already. Any excitement is surely from other implications people think the demo has.

But unfortunately there are no interesting implications from what this site shows. Yes, it demonstrates the boring fact that: "it isn't clear how different people assign different color labels to the same physical stimuli" (and yes, this is FALSELY assuming that everyone's monitors/screens are the same too), but if you didn't already know this... I'm not sure exactly what social context you could have possibly grown up in.

D-Machine··on Is my blue your blue? (2024)
This is the wrong way to do it, psychometrically, see here: https://news.ycombinator.com/item?id=47929056. You need to provide people gradations, or you get junk responses / abandonment, and your instrument doesn't measure what you think.
D-Machine··on Is my blue your blue? (2024)
Wrong way to do it. We know from psychometrics that forced binaries like this just create junk (people disagree with the question, so just choose a forced answer based on some heuristic for each such question like "closest to my mouse / finger" or "most socially desirable" or "same as last time"). So you aren't measuring what you think when you force choice like this.

If you're going to go with linguistic self-report and a single item, you really want something like an 11-point Likert scale. A smart design might get e.g. a person's rating of "blue-ness vs. green-ness" on an 11-point scale, then determine the optimal cutpoint via e.g. clustering, logistic regression, or some other method, to really get something meaningful.

D-Machine··on Is my blue your blue?
Sapir-Worf and its ilk (if we don't have the language/concept, we can't perceive the difference/thing) are widely disproven and debunked, and don't even pass the smell test (learning new concepts and perceiving new things would be impossible). That kind of thinking is so tedious and decades out of date with modern cognitive science, neuroscience, psychology, etc.
D-Machine··on Is my blue your blue?
Agreed, there is no clear premise. Of course that different people looking at the same object will use different colour words is a triviality that anyone over, say, 10 years old knows. If that's the premise of the site, it is boring. People are getting excited because they think this implies something about differences in vision or perception... but it doesn't, that requires much more cleverness to test.
D-Machine··on Is my blue your blue? (2024)
But that is wrong. This doesn't test colour perception or vision, it tests verbal classification of colour perception into a forced binary. Everyone could be perceiving the colour qualia 100% identically, but simply choosing different linguistic cutpoints, meaning you can't say this is about vision / perception at all (it may just be about language use).
D-Machine··on Is my blue your blue?
Asinine and meaningless. Forces a classification on something that obviously anyone with fully-functioning colour vision will classify as "aquamarine" or "turquoise" or etc.

This has nothing at all to do with colour perception, or, if actual differences in perception are involved, this test fails to distinguish those from individual differences in assignment to linguistic categories.

EDIT: To actually test something like this, you need to make an assumption that cannot easily be tested or supported by evidence.

E.g. say we could all agree that, generally, blue + orange is a more pleasant pairing than blue + green. One might then imagine a series of images using orange + varying interpolations between blue and green, with the prompt being "is this combination of colours more or less aesthetically pleasing than the last". The average cutpoint could then be interpreted as a subjective judgement of where e.g. teals become "more blue", from an aesthetic / complementary standpoint. But this test does nothing of the sort.

D-Machine··on The math that explains why bell curves are everywhere
This was sort of my reading as well: I took "clumping" to mean "bump-shaped".
D-Machine··on Ask HN: How do you deal with people who trust LLMs?
Right, I think (hope) the OP meant not to emphasize the "search" in the sentence, but the "reputable source". Of course a Google search now is much worse than an AI search.

And it is the ultimately the reputable source that matters, and whether the person actually read it and checked that the details matched the summary (be it human abstract, LLM-generated, or otherwise).

D-Machine··on Executing programs inside transformers with exponentially faster inference
In the section "Programs into weights & training beyond gradient descent", near the end, they say:

    [...] *the compilation machinery we built for generating those weights** can go further. In principle, arbitrary programs can be compiled directly into the transformer weights, bypassing the need to represent them as token sequences at all. [...] [my emphasis]
In the same section, they also continue:

    Weights become a deployment target: instead of learning software-like behavior, models contain compiled program logic.

    If logic can be compiled into weights, then gradient descent is no longer the only way to modify a model. Weight compilation provides another route for inserting structure, algorithms, and guarantees directly into a network.
So they (almost-invisibly) admit they compile in the weights, but make it clearer this was the whole intention the whole time in later sentences.
D-Machine··on The math that explains why bell curves are everywhere
Yup. And in general more heavy-tailed bumps are in fact better models (assuming normality tends to lead to over-confidence). Really think the universality is strictly mathematical, and actually rare in nature.
D-Machine··on The math that explains why bell curves are everywhere
No, but when you get into the nitty gritty of most things sometimes being influenced by extremely rare things, and also that the convergence rate of the central limit theorem is not universal at all, then much of the utility (and apparent universality) of the CLT starts to evaporate.

In practice when modeling you are almost always better not assuming normality, and you want to test models that allow the possibility of heavy tails. The CLT is an approximation, and modern robust methods or Bayesian methods that don't assume Gaussian priors are almost always better models. But this of course brings into question the very universality of the CLT (i.e. it is natural in math, but not really in nature).

D-Machine··on The math that explains why bell curves are everywhere
This is also right I believe, normal distributions are not ubiquitous really, just they are approximately ubiquitous (and only really if "ignoring rare outliers", and if you also close your eyes to all the things we don't actually understand at all).

The point on convergence rates re: the central limit theorem is also a major point otherwise clever people tend to miss, and which comes up in a lot of modeling contexts. Many things which make sense "in the limit" likely make no sense in real world practical contexts, because the divergence from the infinite limit in real-world sizes is often huge.

EDIT: Also from a modeling standpoint, say e.g. Bayesian, I often care about finding out something like the "range" of possible results for (1) a near-uniform prior, (2), a couple skewed distributions, with the tail in either direction (e.g. some beta distributions), and (3) a symmetric heavy-tailed distribution (e.g. Cauchy). If you have these, anything assuming normality is usually going to be "within" the range of these assumptions, and so is generally not anything I would care about.

Basically, in practical contexts, you care about tails, so assuming they don't meaningfully exist is a non-starter. Looking at non-robust stats of any kind today, without also checking some robust models or stats, just strikes me as crazy.

D-Machine··on The math that explains why bell curves are everywhere
Came here basically looking to see this explanation. Normal dist is [approximately] common when summing lots of things we don't understand, otherwise, it isn't really.
D-Machine··on Can I run AI locally?
These are some really great explicit examples and links, much appreciated.
D-Machine··on Executing programs inside transformers with exponentially faster inference
I'd tend to agree, the only good points I've seen were made by @hedgehog [1] here in this thread:

    I'm not sure about the rest but a significant problem with high frequency tool calling (especially in training) is that it breaks batching.
and then later by @ACCount37 [2]:

    I'm less interested in turning programs into transformers and more interested in turning programs into subnetworks within large language models.
In theory, if you can create a very efficient sub-net to replicate certain tool calls (even if the weights are frozen during any training steps, and manually compiled), this might help with making inference much more efficient at scale. No idea why in general you would want to do this through the clunky transformer architecture though. Just implement a non-trainable, GPU-accelerated layer to do the compute and avoid the tool-call.

[1] https://news.ycombinator.com/item?id=47367986

[2] https://news.ycombinator.com/item?id=47363909

D-Machine··on Executing programs inside transformers with exponentially faster inference
Except their process isn't actually differentiable, as they admit near the end of the post, they just sort hand-wavily suggest that approximately differentiable methods "should" work. Also no mention at all of what the training data would be, where it would come from, or how a loss function could be constructed to continuously score "partially correct" programs (of what that would even mean, or if that idea is even coherent).

What was a good point, mentioned by @hedgehog in this thread (https://news.ycombinator.com/item?id=47367986), is that tool-calls break batching a lot, so there could be huge efficiency gains at scale if you can just pass through a computation sub-network (even if that sub-net is frozen and can't be updated, and is programmed in manually rather than trained in).

Why on Earth you'd want that sub-net to be a clunky transformer rather than just an efficient, GPU-accelerated custom non-trainable layer, though, is unclear to me.

D-Machine··on Executing programs inside transformers with exponentially faster inference
The problem is they are talking about tricks for compiling VMs into transformer weights, which is basically unrelated to actually training transformers on data via gradient descent. Once you get into this actual messy practical reality, you have non-trivial stuff like sparsemax and the Gumbel-Softmax trick to get some desirable improvements to things like the softmax, without all the gradient destruction of things like top-k approaches, but usually at pretty serious other costs (most approaches using Gumbel-Softmax I have read essentially create a bi-level optimization problem that is claimed to be "solved" by some handwavey annealing, but which is clearly highly unstable and hard to tune. I don't know if things have improved here since I last read on it).

So the issue isn't if there aren't ways to effectively approximate their approach, from a strictly numerical approximation standpoint, it is that other factors matter much more in optimization when training on actual data.

D-Machine··on Executing programs inside transformers with exponentially faster inference
If you read the section "Richer attention mechanisms", you can see, no, the mechanism is not generally useable (it requires significant modification to become differentiable). They later speculate:

    While we do not yet know whether exact softmax attention
    can be maintained with the same efficiency, it is easy to
    approximate it with k-sparse softmax attention: retrieve
    the top-k keys and perform the softmax only over those
but if you have played around with training models that use e.g. topk or other hard thresholding operations in e.g. PyTorch (or just think about how many gradients become zero with such an operation) you know that these tend to work only in extremely limited / specific cases, and make training even more finicky than it already is.
D-Machine··on Executing programs inside transformers with exponentially faster inference
One of the worst sentences in the article, clear example of pseudo-profound bullshit, almost certainly LLM-generated.
D-Machine··on Executing programs inside transformers with exponentially faster inference
Nope, they encoded or compiled in a simple VM / WASM interpreter to the transformer weights, there is no training. You'd be forgiven for this misreading, as they deliberately mislead early on that their model is (in principle) trainable, but later admit that their actual model is not actually differentiable, but that a differentiable approximation "should" still work (despite no info about what loss function or training data could allow scoring partially correct / incomplete program outputs).
D-Machine··on Executing programs inside transformers with exponentially faster inference
The model isn't trained, it isn't differentiable (read carefully to the end: they say their model might still work if they made it differentiable, but they don't know), and it isn't clear IMO it could ever be made trainable (what is your loss function that scores a "partially correct" program / compiler, and how are you getting such training data?).

You need non-linearity in self-attention because it encodes feature / embedding similarities / correlations (e.g. self-attention is kernel smoothing) and/or multiplicative interactions, it has nothing to do with determinism/indeterminism. Also, LLMs are not really nondeterministic in any serious way, that all just comes from tweaks and optimizations that are not at all core to the architecture.

D-Machine··on Executing programs inside transformers with exponentially faster inference
LLMs (or at least transformer-based LLMs) are effectively almost entirely deterministic, the randomness being largely only present due to (unnecessary) optimizations and other tweaks.

Temperature is not at all core to LLMs, it is something that rather makes the outputs more varied and desirable for human consumption generally. It is trivial to set to zero for applications like this.

On CPUs, the models are essentially fully deterministic, even with FP accuracy, and most common kernels have reproducible (albeit slower) variants even on GPUs. Otherwise, yes, FP non-associativity on GPUs is the only real source of randomness in inference.

The other issue arises from batch invariance, but this is a problem that occurs only at scale when serving multiple users / inputs have some randomness too. You can (usually) trivially eliminate this by controlling what goes in the batch or making the batch size be one. There are also other more clever mitigations for this, none of which are secrets.

EDIT - Forgot reference: https://thinkingmachines.ai/blog/defeating-nondeterminism-in...

D-Machine··on Executing programs inside transformers with exponentially faster inference
There is no training in the usual sense of the term, i.e. no gradient descent, no differentiable loss function. They use deceptive language early on to make it sound this way, but near the end make it clear their model as is isn't actually differentiable, and in theory might still work if made differentiable. But they don't actually know.

But IMO this is BS because I don't know how one would get or generate training data, or how one would define a continuous loss function that scores partially-correct / plausible outputs (e.g. is a "partially correct" program / algorithm / code even coherent, conceptually).

D-Machine··on Executing programs inside transformers with exponentially faster inference
Here is a rough list, some may be contentious individually, but the more of these appear, the more you should suspect an LLM:

Cadence and rhythm: LLMs produce sentences with an extremely low variability in the number of clauses. Normal people run on from time to time, (bracket in lots of asides), or otherwise vary their cadence and rhythm within clauses more than LLMs tend to.

Section headings that are intended to be "cute" and "snappy" or "impactful" rather than technically correct or compact: this is especially a tell when the cuteness/impactfulness is deeply mismatched with the seriousness or technical depth of the subject matter.

Horrible trite analogies that show no actual real understanding of the actual logical, mathematical, or visuo-spatial relationships involved. I.e. analogies are based on linguistic semantics, and not e.g. mathematical isomorphism or core dynamics. "Humans cannot fly. Building airplanes does not change that; it only means we built a machine that flies for us". Can't imagine a more retarded and useless analogy for something as complex as the article topic.

Verbose repetition: The article defines two workarounds: "tool use" and "agentic" orchestration, then defines them, then in the paragraph immediately following, says the exact same thing. There are basically multiple (small paragraphs) that all say nothing at all more than the sentence "LLMs do not reliably perform long, exact computations on their own, so in practice we often delegate the execution to external tools or orchestration systems".

Pseudo-profound bullshit: (https://doi.org/10.1017/S1930297500006999). E.g. "A system that cannot compute cannot truly internalize what computation is." There is thankfully not too much of this in the article, and it appears mostly early on.

Missing key / basic logic (or failing to mention such points clearly) when this would be strongly expected by any serious practitioner or expert: E.g. in this article, we should have seen some simple nice centered LaTeX showing the scaled dot-product self attention equation, and then some simple notation to represent the `.chunk` call, and subsequent linear projection, something like H = [H1 | H2], or etc., I shouldn't have to squint at two small lines of PyTorch code to find this. It should be clear immediately this model is not trained, and this is essentially just compiling a VM into a Transformer, and not revealed more clearly only at the end.

D-Machine··on Executing programs inside transformers with exponentially faster inference
This is my interpretation as well.

EDIT: Actually, they do make this clear(ish) at the very end of the article, technically. But there is a huge amount of vagueness and IMO outright misleading / deliberately deceptive stuff early on (e.g. about potential differentiability of their approach, even though they admit later they aren't sure if the differentiable approach can actually work for what they are doing). It is hard to tell what they are actually claiming unless you read this autistically / like a lawyer, but that's likely due to a lack of human editing and too much AI assistance.

← PreviousPage 6 of 20Next →