652 karma · joined September 18, 2025
The prime example is "this says a lot about society"
...I don't think this will have a positive impact on the world
It's always "yudkowsky is a weird guy" or "they had an orgie once" but who gives a fuck?
I just want someone to lay out, in impersonal terms, the flaws of their reasoning and why I shouldn't agree with them.
My problem is with the whole genre of article rather than this one specifically.
Curious because I'm considering joining a similarly small lab.
The real problem is that when using fewer sampling steps than output tokens, diffusion formulations fundamentally cannot represent distributions where output tokens are heavily codependent. Autoregressive formulations don't have this problem, they can represent any distribution (ignoring limitations of the underlying model).
That would explain a lot of the terrible AI/LLM takes online.
IWAE: sample a bunch of latents, weight the loss of training the model using each one by softmax(-error). For images and text where the errors have large variance, those weights become one-hot, yielding this algorithm.
1. It's hard to measure (and people can disagree about it)
2. It can't really be improved using RL without a human in the loop (which is how math is being trained)
It often feels like the model is ignoring my inputs and just doing what it would expect the bot to do (which is unsurprising if the model could predict what would happen next during training without paying attention to the inputs)
> We see it with self-driving cars. It's really impressive what is possible. But they aren't true self-driving. A person needs to be sitting at the wheel, keeping their attention focussed on traffic as if they would be driving themselves, so they can intervene when the AI makes an inevitable mistake.
There are self-driving cars all over San Fransisco transporting people on public roads with no human at the wheel. This proves my point: those cars are not perfect, but they are human-level (or close), and that's all that's needed.
If you want proof just look at the benchmarks. Modern frontier models can get basically perfect accuracy on American Invitational Mathematics Examination tests: https://matharena.ai/?comp=aime--aime_2026
If you want an explanation of how they do math, we've found geometric calculators inside their neural networks: https://www.goodfire.ai/research/a-geometric-calculator#
If Opus gets all but the hardest questions right, it might have a higher hallucination rate because the questions it gets wrong are the questions where verification or hallucination detection are the most difficult
The first idea (as I understand it as retrieving token ids rather than hidden states) is going to really struggle to do useful compositional reasoning and contextual recall.
The second idea has been been done a million times, with Linear Attention being maybe the first modern example. Hyena, state-space models, DeltaNet, and LaCT also lie in different regions of the performance-parallelizability spectrum of fixed-size models.
This is plainly not true anymore
Their 1.2B model was trained on only 10B tokens, which is less than half of the chinchilla compute optimal number. Modern overtrained 1B LLMs are trained on the order of 10T tokens (1000x more).
This is important because, from my own experience, simplifications and alternatives to standard attention can look fine in the under-trained regime but lag after over-training. This happens because attention has very little out-of-the-gate inductive bias, so it takes a lot of training for the expressiveness to really shine through.
I can't fault the authors since longer training runs cost money, but it warrants pointing out.
I'm also disappointed that they didn't report reasoning benchmark results for the Q=K-V case, since that is by far the most theoretically interesting case (in my eyes).
Do you have an alternative idea in mind?
The model takes in the context, encodes it into a "memory" (the KV cache), and accesses that memory later. That fact doesn't change just because the KV cache grows in size with the context.
I don't know what memory would look like other than an encode-retrieve loop.
Relevant: Transformers are Multi-State RNNs - https://arxiv.org/abs/2401.06104
The entire file is not changed, but the KV cache is.
> It doesn't remember anything
The model definitely remembers previous exchanges within the same conversation.
gzip can be used as a (not very good) LLM-like text and image generator: https://arxiv.org/abs/2309.10668
That is to say, we don't know why they give the outputs that they do.
If we did know how they worked, AI interpretability would not be an open and growing field.
A bold claim given that the current top post on HN is "An OpenAI model has disproved a central conjecture in discrete geometry": https://news.ycombinator.com/item?id=48212493