Building Transformers from Neurons and Astrocytes
biorxiv.org
biorxiv.org
To clear up some confusion:
1) The paper proves an architectural equivalence between two models: neuron-astrocyte networks and Transformers. This is different from studies of representational similarity, which look at Transformer "responses" and compare them to brain responses using regression techniques.
2) The paper provides testable hypotheses for what astrocytes are doing. We can either succeed or fail at rejecting those hypotheses after comparing to real brain data. That is future work.
Happy to answer questions.
- Leo
f(x; W, W', ...) = ...max(0, W''max(0, W'max(0, Wx)))
And called it a "neural network"When your opening gambit is that ridiculous, you really can't expect them to improve. We're not dealing with people interested in science in the ordinary sense.
I think we're in an age of pseudoscience where rather than wear lab coats, pseudoscientists "wear statistics" in a self-deluding attempt to ape science.
However, I think this paper falls closer on the "cooky speculation" side. It does, nevertheless, participate in the culture of pseudoscience around these issues.
It would be great if the paper owned up to its radical speculative nature, and didnt brush aside it's lack of parameterisation on time (which is pretty fatal to taking this seriously). Then it would be a pretty good model for what "radical considered speculation" looks like which we could contrast with the genuine pseudoscience of many papers in this area.
Why not explain the point that you disagree with and explain why you disagree? That would be valuable for the conversation.
The reason it's, at best, speculative is that we have no idea at what layer the brain-body-env performs computations.
If we're interested in a dynamical model of the brain-body-env, then we want a series of partial differential equations which are parameterised by space and time. Just as we would describe the electrical field of a digital machine.
However since we know that +/- voltage states of the machines we create are the relevant discrete states to give a computational description, we can also give a non-dynamical account of a digital computer (ie., we can say it realises the sequence: 0001, 1001, 0101, etc.
What this paper does it take arbitrary aspects of arbitrary parts of the brain, and tries to marry them to some existing computational framework.
If an alien were to analyse a CPU that way, for example, it could easily conclude that it's the heat of air particles warmed by the CPU which "did computational work". Since they do computational work, but none relevant to the digital computer we are using.
This paper treads the line, some others cross it, in breathlessly claiming their series of wild, indefensible, and pretty clearly wrong assumptions which motivate their fashionable computational framework (ie., NNs, etc.) show that their press-released new projects (eg., GPT, etc.) are like brains somehow.
Can you try describing what they specifically claim that you find arbitrary? Like, what are the “clearly wrong assumptions?”
Even if we buy that parts of the brain have some properties which discretely encoded are "plausibly modelled" by a transformer network, this claim isnt surprising. It isnt surprising applied to a plant.
In other words a descriptive "computationalist" project is almost always going to succeed. I can describe the orbit of the earth around the sun as 010101...
But the transition from the earth in front of the sun (0) to it being behind (1) simply describes what it is we're trying to explain.
Likewise pretty much every possible system has some aspect that can be described by the sequence 0101010...
`F = GMm/r^2` however is not a descriptive project. It provides properties, causes, measures, and so on... it says what exists, what it is like, how it functions, and how the target of analysis (gravity) is controllable... ie., we can produce more of it with more mass, or less distance.
For computational explanations to count as explanations they must elucidate some explanatory aspect of the system's function.
In general, it's very hard for computational descriptions to do this; and this issue is made worse because the people who use it are (discrete) mathematicians and not scientists and really think that saying "orbits are ellipses", "leaves are ellipses", so "orbits are like leaves" is an actual explanation.
The purpose of a transformer network, as used by data scientists, is not captured by brain function, nor by plant function. And finding a pattern in the brain which could be modelled this way, carries over no "explanatory power" of transformers as ordinarily used.
This mistake is at the heart of much "computationalist pseudoscience".
> System S has abstract discrete properties P, System T so too has them, so our explanations of T provide explanations of S.
This is magical thinking, and a product of the people in this game never having really been familiarised with experimentation and the scientific method.
Are you saying it isn’t enough to have a model? That it must be accurate in all ways to be an adequate explanatory model?
Describing the spin of an atom as spin is explanatory—but only incompletely. It still has value.
In this case, the article described how astrocytes and neurons interact in some way that can produce properties of transformers. Correct? Why, in your opinion, is this a fruitless end? We won’t find a memory system or CPU either, but we may find accumulators and registers, and we are damn likely to find convolutions.
Why not attempt to explicate a mechanism for transformer-based computation in the brain? Maybe they are wrong. But to call this pseudoscience…
Attributing spin to particles is explanatory. You can, for example, perform the stern-gerlash experiment.
The problem is something like this, take the formula f(a, b, c; A, B, C) = a^A.b^B.c^C
That formulae describes gravity (a=m,b=M,c=r; units of GN), electrostatics (a=q,...), pendulums, oceans, thermodynamics... etc.
The formulae is a pure syntax on to which we can impose a wide variety of radically different semantic. Knowing that formulae can be used with a system is meaningless. It can be used to describe an aspect of almost any system.
The explanation part comes from the semantics of that formulae.
What computational descriptions do is simply show that some formulae can be used. A priori, we basically already know that even, eg., that formulae above "describes the brain". That doesnt mean the brain is anything like gravity, thermo, etc.
The false equivocation here is this: since the syntax of transformer models applies to the brain, their semantics as used on digital computers shows the brain is like a digital computer.
This is a false equivocation, and extremely unlikely.
If people want to make the above claim they need to perform an experiment which shows that the brain is operating like a digital computer, that a formulae applies to both is unsurprising.
We should expect some aspect of the airflow in a room to be describale with the syntax of a transformer.