The illusion of state in state-space models
arxiv.org
arxiv.org
The fact that we have already obviously developed a technology which has an excellent grasp of language, can select between tools, perform knowledge retrieval from external databases, perform some form of simplistic reasoning and planning… it really feels (again, only intuitively!) like everything outside the perimeter of LLMs is where the actual frontier of AGI will be.
It feels like with enough sensorimotor affordances and mechanisms for saving and accessing memory, all that is left is the (obviously complex) framework for connecting all of these together, with potentially many instances of each module working together… especially when you consider that one such system can potentially use human actors as tools (eg “hello fellow human, please complete this captcha for me, as I am low-vision”).
I guess more succinctly - I do not believe that GPT-4 et al are remotely close to AGI on their own, but I do feel (again, this is just “vibes,” and obviously people smarter than me disagree!) that GPT-4 could be a component combined with lots of other tech we already have to essentially achieve AGI already. Perhaps what I am describing would just be a pale imitation of what others mean by AGI and be reductively called another mechanical Turk. I am probably the misguided one, but nonetheless, I can’t escape the feeling that “the one model to rule them all” is a red herring vis-a-vis a complex network (ha) of modules interacting together to achieve meaningful agency. Maybe it’s just that I am more interested in “meaningful agency” as a guiding principle over “AGI.”
Text generation can only get so much better but the other modules required to simulate intelligence have tons of room for improvement
The paper isn't making that argument. The point it makes is that parallelisable SSMs are theoretically no more powerful than transformers, contrary to some people's assumption that they'd be theoretically equivalent to RNNs and hence more efficient at certain kinds of problems.
Gates and non-linearities in RNNs allow the state to be "any" function of previous inputs.
I don’t think it will hold us back. If anything, it’s very exciting to see how many people in the ML field are challenging the status quo from many, different angles.
These proofs still hold; pure MLPs (without a modern activation function) aren't very useful, both in theory and practice. What made them useful was the realisation that combining them with a proper activation function makes them much more useful, both theoretically and practically; this discovery took time.
XOR is simply not linearly serperatable, requiring an MLP or kernel trick still holds.
It is a similar reason that attention works for majority gates but not parity gates in the general case.
Acknowledging that reality resulted in new developments, but is still a limitation.
Perceptrons are binary classifiers.
From my reading, "non gated SSMs" are LTI systems. They're linear (the paper ignores the activation function) and the "does not depend on input" criteria is the controls way of saying "time invariant" (replace time with the domain of the input vector). It is not surprising that a system that cannot adapt to changing input conditions (by construction) is fundamentally limited to the problems it can model. This is why adaptive filters are studied - LTI systems are useful computationally, but limited because they cannot adapt. Relaxing the "TI" constraint greatly expands the domains of problems that can be solved.
When the feedback matrix is diagonal you have an FIR system. While "finite impulse response" has as specific mathematical definition, conceptually it means the state at time n depends solely on past inputs and states n - N where N is the size of the state vector. So of course, if the feedback matrix is diagonal, the system is limited in the kinds of stateful problems it can handle.
What this paper is missing is that connection to the limits of LTI and connection to TC0, which is very interesting indeed.
In other words, the state of the network is still infinitely long and definitely not an illusion. However if the problem requires the network to adapt to its input then a non-gated SSM is not going to be sufficient. That's an interesting research space.
The field completely appropriated and redefined so many terms of common language, that now it’s hard to talk in plain language with someone formally trained in physics about physical phenomena
For example, everyone has some sort of intuitive idea about what Energy is. But if you use that word with a physicist, watch out, for them it means something super specific within the context of assumptions and mathematical models and they will assume you don’t know what you are talking about because you are not using their definitions from their models
Same thing happens with infinite in math
Apologies if you know the following already, but maybe others reading your comment feeling similarly will not be familiar and might be interested.
At least intuitively, I like to motivate it this way- pick your favorite simple state space problem. Say a coupled spring system of two masses, maybe with some driving forces. Set it up. Perturb it. Make a bunch of observations at various points in time. Now use your observations to figure out the state space matrices.
There’s fundamentally not really anything different (in my opinion) with using Mamba (or another state space model) as a function approximation of whatever phenomenon you are interested in. Okay Mamba has more moving parts, but the core idea is the same: you are saying that on some level, a state space is an appropriate prior for approximation of the dynamics of the quantities of interest. It turns out being pretty remarkable the number of things this can work out quite well for. For instance, I use it to model the 15-min interval data for heating, cooling, and electricity usage of a whole building given 15 min weather data, occupancy schedules, and descriptions of the building characteristics (eg building envelope construction, equipment types, number of occupants, etc).
h = Ah + Bx
y = Ch + Dx
where x is the input, y is the output, and h is the state. They use "h" instead of "s" for the state variables because they're called "hidden states" in the literature. edit: it is obnoxious they've flipped the convention for A/B/C/D which is the one thing controls people agree on (we can't even agree on the signs and naming of transfer function coefficients!).Where this diverges from dynamical systems/controls is that they're proving that when x/h/y are represented with finite precision numbers, the model is limited in the problems it can represent (no surprise from controls perspective), and they prove this by using an equivalence to the state-space formulae that's consistent with evaluating it on massively parallel hardware.
The classical controls theory is not super applicable here, because what controls people care about (is the system stable, is its rise/fall time in bounds, what about overshoot, etc) is not what ML researchers care about (what classes of AI problems can be modeled and evaluated using this computational architecture).
Do Microsoft & friends who are about to build trillion dollar AI data centers know about these proven limitations of transformer-based architectures?
The only open question is whether people have common-enough queries for this charade to work out. It seems there's quite a lot at least. But this number will decrease over time for various reasons. So it's a game of building a system that can be retrained on the answers people are giving it fast enough that people don't notice where the answers are coming from.
Statistical AI is just a way of sampling from a historical dataset with a similarity metric. It only works to answer questions if you're sampling from a (Q, A) database in the same language the user already understands. The question was answered by a reasoner, it is now answered by a system which replays answers.
(I have no idea if it can! But I'm confident enough that I'm willing to just say it can, at risk of being proven wrong. If GPT can't do that, I'm fundamentally misunderstanding how it works - that is, I don't have a paper offhand showing that GPT shares concept neurons between languages, but I'm willing to bet I could find one if I went looking.)
In other words, if you co-trained GPT-4 on Earth Internet and Alien Internet, there's a good chance it'd end up able to translate English to Alienese, purely as an emergent ability, if it had the concept of translation at all.
Intelligence is compression. With sufficient abstraction (layers) and sufficient volume (dataset), any description of the same reality will assume the same structure. And no learning algo worth its salt will keep two identical structures around.
Likewise, languages do not have the same distributional structure. There is no reason an alien language would use discrete tokens to name properties, nor co-locate tokens by linear position in a 'sentence'. Historical human languages did not; using, eg., the full 2D structure of the clay tablet.
To suppose that it is the glyphs and their colocation which somehow bare meaning is a nonesensical superstition. The world is what our words mean, and it is we, in that world, who provide them their meaning. We can do so with arbitrary linguistic structures.
By analogising to any animal or human mind, you aren't describing anythign that acutally exists. A neutral network algorithm isnt neural and it isnt a network. It's a statistical curve-fitting algorith. One oughtnt study trees to understand a decision tree either.
This language is entirely metaphorical. There are no neurones in an NN, there are just summation entires in a matrix. This matrix comprises weights, which define the orientation and scale of the line pieces which form the curve being fit to the data.
We're not talking about 2-layer perceptrons anymore here.
In my opinion, you should look less at the formalism and more at the empirically demonstrated performance.
Whatever property you might imagine a transformer architecture to have (and it is vastly fewer than the set needed for general computation), the problem here is that it's being applied to approximate the structure of historical text data which isnt being generated from such a function.
Indeed, there is no function which generates text data, there's a very large number of independent generating processes that give rise to the distribution of text. The phrase "the war in ukraine" acquires a different semantics over 2010-2030 in a radically different way than, "I liked that film" does.
The capacities which produce distributions of text tokens are highly varied, complex, invovle a vast array of our mental processes, and so on. There's literally almost nothing in the distribution of text tokens that corresponds to any features of these processes.
The structure of language is conventional, and rests on our familiarity with such conventions. Otherwise, let's end all science, everything to be known about the world derives from how "e" occurs alongside "lectron"
Our mental capacities can be patterns. https://en.m.wikipedia.org/wiki/Predictive_coding
> There's literally almost nothing in the distribution of text tokens that corresponds to any features of these processes.
There's nothing in evolutionary fitness that necessitates intelligence or reasoning ability either. Yet here we are.
Sorry but if you don't see the connections then you need to do some reading on theory of mind, cognition, information theory, physics, philosophy. All of the fundamental basis are met to allow reasoning to emerge in LLMs.
It is clear now where your confusion lies, and why you are led to believe so strongly that LLMs cannot reason: it is because you are an ML practitioner you overweight your expertise yet you don't know what you don't know, and have foundational gaps in your knowledge. You lack the context in these other fields. If you had them, your position should be closer to agnostic than this strong belief of yours that LLMs in their current form cannot reason.
One does not form "connections" between them as in some wide-eye conspiracy theorist.. "predictive coding" has little to do with "prediction" in the ML sense. The latter concerned with making a quantitive estimate of some variable by summarising historical data.
What we are doing when we revise a "mental model" is done by counter-factual simulation of possible future states. Statistical AI models conditional probability structures, and computes predictions as an expectation over weighted summarised historical data. This is not a means of performing counter-factual simulation.
One trivial, sadly empirical, way to see this is to note that each marginal token generated is of constant time and energy use. Yet trivially, reasoning and a variety of other mental capacities should require abitarily different time to run. Eg., simulating a complex scenario is necessarily more intensive than a simple one, and so on.
Yet I am quite annoyed that we need such dumb observations to make this point. It speaks of a profound ignorance of "theory of mind, cognition, information theory, physics, philosophy" and especially neurology and zoology which are your most significant missing terms.
It is no real mystery what the structure of various mental capacities involves; nor any mystery what s(Ws(Ws(WX+B)+B...)...)...) computes. Even involving anything beyond trivial applied statistics and trivial results in science shouldnt be required here. This stuff is very obvious.
But those semantics are revealed in the greater context of the phrase! That's why it's so important that transformers can attend to large context ranges; that's what lets them learn the greater semantic patterns to begin with. And at the limit, at a scale greater than phrases, I simply reject the idea that the same article - the same book - can have totally different meanings depending on context. Language isn't just shaped by context, it shapes context itself. Because language cannot be considered without context, it reveals information about that context, and in fact any compression of language ultimately requires modelling the person and even society that produced it. That's what the network learns.
If your words don't have meaning beyond themselves, what are you even talking about?
...err... of course? That's the whole point.
"Context" here isnt other words. It's the world. The meaning of words is the world.
When I say, "I like what you're wearing" i'm not summarising a history of prior texts; it has nothnig to do with any statistical operation over historical documents. It has entirely to do with what you're wearing.
Langauge use is a side-effect of being embedded in reality, directly attentive to it, and so on. Words are mere symptoms of how we are situated in the world.
LLMs merely replay these back to us. They are not in the world. They cannot, in principle, ever mean, "I like what you're wearing"
The strings 010101010111100101, 000010101010111010101, 120192019092109209102, etc. are all "about" the exactly the same scenario whereby different communities of speakers adopt different conventions for the meaning of the terms (indeed, whether pairs/tripples/etc. are terms).
This is not like a thermometer in a glass of water, the height of the mecury being determined by the temperature. Nothing about language is determined by reality. It is purely a highly lossly encoding scheme based on historically arbitary encoding conventions. It isn't a measure of reality in any sense.
Modelling this encoding scheme doesn't model anything about the world. It only appears to if you generate text according to this scheme.. and only to those who can decode it. Hieroglyphs likewise, scratches on clay, and everything else.
You are subject to the illusion of intentionality whereby your phenomenal experience of langauge is rendered immediately meaningful by your brain, ie., your decoding scheme is prior to the construction of your experience of the world. This is wholey absent from anything machines are processing.. whih are only the frequency relationships between tokenizations of syntax
This is just transparently silly. Or rather, you're equivocating between "language is not fully determined by reality" and "language is uncorrelated with reality." Language is partially caused by reality; thus, large language models, seeking compression, learn the arbitrary parts and the reality parts separately. That's why they can translate.
The human sensory stream is a kind of language. The exact same logic goes for it! If LLMs could not learn about reality from language because language was not determined by reality, then human brains could not learn anything about reality either.
Nothing about reality forces a certain neural state vector to correspond to a certain sense impression. E pur si muove.
As in you think you are learning a speed limit for a road by looking at a sign with a speed limit on it. But the sign itself is just some shapes, and no aspect of its structure has anything to do with speed, roads, or speed limits. Literally, everything here is arbitrary convention; and under arbitrary permutation, the sign can mean the same, and share no properties with any other permutation whatsoever.
Rather it is by being-in-the-world you have acquired a direct understanding of roads, speeds, limits, rules, laws, society, etc. And alongside this all the while internalizing arbitary conventions by which we signify these things.
A stop sign could look like anything at all. We just decide to use arabic numerals, red, white, a certain eye-hight, etc.
The meaning of the sign is something you are reconstructing prior even to being cognitively aware of the sign you're looking at. The meaning of these conventions is a mental reconstruction made possible by your direct familiarity with the world these conventions are related to.
You first act within the world by building a rich conceptual understanding, you then associate arbitrary conventions with this understanding, and then can make inferences about this world by reconstructing a novel conceptual understanding given by these signs.
There is literally nothing at all in speed signs that has anything to do with their meaning.
This isnt a speculative observation. It places a very hard limit on what the entire field of statistical AI can do, let alone LLMs. Any model created by empirical risk minimization, or fitting-to-historical-data, lacks any relevant capacities. It has only one: drawing a data point from some inferred distribution over historical cases.
In any case, this isnt a speculative exercise. Statistical AI will be forever "limited" to a fixed computational budget for computing predictions over historical cases, with necessarily limitations this involves -- this is only an engineering problem where these cases aren't representative of the kinds of queries people have.
eg., they will fail under relevant counter-factual permutation of the conventions of the language/domain/topic, etc. -- see, e.g., https://arxiv.org/pdf/2307.02477
Applying the lessons of the above paper, note chatgpt3.5 gets the evaluation of python wrong under permuation of its syntax: https://chatgpt.com/share/7df48393-9cb8-41f4-a84e-d21b680702...
..whereas ChatGptv4o does not. This suggests to me that there's been a very large amount of prompting by users since 3.5 trying to create novel programming languges in this way.. hence 4o, having been trained on subsequent user prompts, its "better" at this task.
Of course nothing important has changed between v3.5 and v4o in terms of what the model is, rather there is more query-relevant data to sample from.
I just don't think this is correct. Language is embedded in reality, is a product of reality, and is correlated with reality.
> As in you think you are learning a speed limit for a road by looking at a sign with a speed limit on it. But the sign itself is just some shapes, and no aspect of its structure has anything to do with speed, roads, or speed limits.
Again, simply wrong. 5km/h limit, 60km/h limit, 120km/h limit - even the physical width of the text has a correlation to the exponent of the speed limit! Which will correlate to the curvature of the road! These are not arbitrary, they're tightly interlinked on every level except the specific shape of the glyphs themselves. That's why your argument only holds up for the most basic features such as glyph frequency. As you zoom out, structural similarities begin to appear; not in the shape of the text but in the shape of features abstracted from the text. That's why LLMs attain novel capabilities as you increase depth and context size, as they can begin to condition on the deeper structures of reality that leak through the language. As language is abstracted from reality, at sufficient scale, a cognition based on language begins to necessarily model reality itself.
> You first act within the world by building a rich conceptual understanding, you then associate arbitrary conventions with this understanding, and then can make inferences about this world by reconstructing a novel conceptual understanding given by these signs.
Completely unclear if this is true at any rate. Humans don't first learn to interact with the world and then learn language; language and world-understanding co-develop.
(Though not quite in the same sense.)
In principle languages are arbitrary, but in practice they’re describing the same world and the same concepts end up being useful.
Yes, if you construct an embedding vector on texts-A, and another on texts-B where (A, B) are translations of each other, then an "unsupervised" algorithm really will give you the dizzying heights of a little above coin-flip accuracy on highly engineered self-selected benchmarks.
They evaluate by by looking at the most in-use words in each vocab.. so you take the most in-use words on translations of Wikipedia, whose frequency is decided by the need of translation.. and then you use that to evaluate.
It is blindingly obvious that the structure and frequency of heirglphys on tombs, Chinese glyphs in poetry, and latin in medieval liturgical literature are not distributed by Reality.. written in this order by God so that the Langauge of Reality is what places "d" alongside "oor". We already know this to be the case. The assumption of its opposite is rank pseudoscience.
I mean, on the first level, of course they're entirely determined by reality in the sense that the human brain is a real, physical object. But also on a second level they're still entirely determined by reality because the shape of the human mind is also entirely determined by reality. What is the mind for except reflecting reality? What is language for except communicating it? Sure that reality is warped, filtered, reduced and biased, but the data is still in there. That's in large part why LLMs need such ludicrously large training runs.
I don't think letter frequency is objectively determined, but we know for a fact (many studies!) that the features that large language models learn are far, far above the scale of letters. Even arguing about phrases isn't engaging with the current state of the art.
We're not talking about Markov chains here.
There are two hypotheses: H1, the structure of a response from any given prompt is computed using distributional properties of historical data; H2: the response is computed via deduction from premises to conclusions of agent employing the semantics of the terms, their logical connections, and connections of relevance.
In many cases a prompt/reply will confirm both hypotheses, hence confirmation bias and why we dont bother "confirming" any hypothesis. Rather to choose betweeen them, if you wanted to use data, you just find cases where reasoning fails in such a way that H1 is the more plausible answer. Such cases are easy to find, and across the literature.
This whole thing is pointless however, because it's blindingly obvious from what a statistical AI algorithm does, which is empirical function fiting on historical datasets to approximate conditional probability distributions. This process is extremely well-understood, and we know necessarily that it is just approximating a historical distribution.
Prove that this precludes formation of mind. Prove that human minds do not work this way.
Until such premises are established, your argument is not sound. This leaves open the possibility what we're seeing is in fact a degree of reasoning ability in LLMs.
If you reflect on that carefully, you'll note the only way this is possible is if the algorithm does not use any semantic features of the input. Indeed, if it has no relevant capacities at all.
Consider e.g., the prompt, "imagine a world where..., and then infer..., and then what would be... ?" for various "..." we can choose, the output ought take arbitrarily longer to compute; but its constant.
For any word: imagine, reason, suppose, simulate, believe, remember... denoting any alleged mental operation, we can trivially construct cases where actually performing this operation would take more or less time.
"Given premises A, B, C, D,...; infer conclusion..." has a prompt reply which is constant in time regardless of how many premises we input. "Recall memories A, B,C,.." likewise.
The only way that answering "yes" to a question, say, regardless of what that question is, taking the same time/energy/etc. to answer.. would be if no semantic feature of that question was being used in its answering.
This is an LLM, it's all statistical AI: the number of operations performed per prediction is constant for all predictions. Almost no alleged capacities are consistent with this fact.
This proves humans cannot think.
The point is that if I ask you to do some complex task that requires a one-word answer, you will take a long time to think about it. If I ask an LLM, it will always take the same amount of time, regardless of the task.
It follows from P =/= NP that this should necessarily not happen if the LLM is actually evaluating the semantics of the question. However, it also follows from just the properties of what capacities are being claimed
Reasoning, as an algorithm, is at least, O(NumberOfPremises) etc. The number of operations performed, in every case, for all statistical AI algs, is the same. It's not O(ANY PROPERTY OF THE PROBLEM) -- this is a catastrophe for anyone who would claim that a statistical algorithm considers the problem at all.
Statistics here is just a very naive short-cut. Rather than solve the problem, you're just sampling from previous answers.
This, of course, should be obvious to any one with the bare minimum of knowledge of applied statistics and the formalism of ML. Neverthless, of course, it isnt. One has to imagine that overwhelming amounts of corporate propaganda and scifi wishfulfilment are at work.
You note this is offtopic, but I disagree. This is part of your fundamental misunderstanding of how to make good use of LLMs - and how LLMs make good use of context space.
> The point is that if I ask you to do some complex task that requires a one-word answer, you will take a long time to think about it. If I ask an LLM, it will always take the same amount of time, regardless of the task.
This is simply not the case. I will take a long time to think about it, because my brain is generating tokens in the conscious workspace. I will take roughly the same time for every token, same as a LLM. (I believe that's what brain rhythms are about.) However, as opposed to a LLM, I can do computation without echoing it to stdout. If you stuck a probe in my conscious workspace, I believe you would find it looking suspiciously similar to the output stream of a multimodal LLM! And there are several papers to the extent of - hey, LLMs do in fact give better answers if you let them insert filler words, or even dots or spaces or invisible tokens - because they learn to use the additional time you give them to -- think.
Cough just like us cough.
That's why all these tasks that tell the LLM to "not output any extraneous information but only give the answer" are so dumb and misleading. They're literally telling the LLM to blurt out the first thing that comes to mind.
In other words, you think the context window is speech, which is why you are confused. The context window is the LLM's axis of time. That their time is conflated with speech is the primary reason why LLMs have a reputation for being glib, surface speakers - and why they continue to underperform. They literally cannot stop to think - not because of any fundamental inadequacy of the technology, but just because that's how we've set them up. I suspect the amount of boilerplate text given before an answer correlates directly with quality of the answer. That's why the energy is misleading - the computational work isn't just the token, but every token between the point where the answer could be computed and where the LLM was forced to commit to a final answer.
And as one could expect from this, if you give the LLM time to think without speaking, their performance shoots up significantly. The moment somebody at OpenAI reads the QuietSTaR paper (GPT-2 scale network with GPT-3 quality!) and understands what it means, the timer to AGI begins.
I think parametric memory and accelerated grokking are equally promising. The convergence of all of these will dramatically improve reasoning. The next few years will be interesting indeed.
An LLM would have the wherewithal to consider the possibility they're wrong, or to be agnostic in their beliefs until we ourselves understand better what it means to reason.