The brain ‘rotates’ memories to save them from new sensations
quantamagazine.org
quantamagazine.org
Now, you want to use this sparse array to represent a note in a song. So you need every note to consistently map to a distinct* sparse array.
However, you also want to be able distinguish a note as being in one song or another. The representation should tell you not only that this is note A but note A in song X.
How might you do that? Well some portion of the ON bits could be held consistent for every A note and some could be used to represent specific contexts.
Stable and variable bits of you will.
Now if you look at two representations of the note A from two songs you'll see they're different. How different are they? Well you could just count the bits they have in common or not, or you can treat them as vectors. (Lines in high dimensional space) Then you can calculate the angle between those two lines. As that angle increases its easier to distinguish the two lines. They won't ever get to full "right angles" between them because of the shared stable bits, but they can be more or less orthogonal.
That's what's happening here. The brain is encoding notes in a way that it can both recognize A, but also recall it in different contexts.
*But not perfectly consistent, we use sparse representations because the brain is noisy and it's more energy efficient. Pretty close is good enough in the brain and you can encode a lot of values in 1000 choose 20 options.
Same goes for what's being alleged here: Is there even a way to visualize this that makes mathematical sense? What will be the corollaries to this discovery simply as a result of what the mathematics of rotations will dictate?
-- Jonathan Richard Shewchuk, from An Introduction to the Conjugate Gradient Method Without the Agonizing Pain
From what I understand, you are saying this rotation is non-intuitive. Could you elaborate more or share some relevant links?
Imagine you have a population of 100 neurons. To each one you associate a real number representing its activity. So a particular snapshot of firing would have 100 real values. It's a vector in a 100 dimensional space. Each individual neuron is one dimension. The high dimensionality of the space is what makes it unintuitive. But, after some familiarization you learn heuristics for reasoning about high-dimensional spaces and combine that with your 3D spatial intuition.
The 45 degree rotation is just some nice art, but you could think of it as representing a projection down to 2D.
I understand higher dimensional connections in theory (such as in an abstract representation of neurons within a computer), but I can’t imagine how more highly-connected neurons could all physically fit together in meat space.
[0]: https://stevenson.lab.uconn.edu/scaling/ [1]: https://www.nature.com/articles/nn.3776 [2]: https://doi.org/10.1016/j.conb.2015.04.003 [3]: https://doi.org/10.1016/j.conb.2019.02.002 [4]: https://arxiv.org/abs/2104.00145 [5]: https://doi.org/10.1016/j.neuron.2017.05.025
[0]: https://www.biorxiv.org/content/10.1101/214262v1.abstract
This isn't really something about neurons per se, it's about systems.
Suppose I have a system that can be fully characterized (for my purposes) by two number: temperature and pressure. If I take every possible temperature and every possible pressure, these form a vector space. But notice that temperature and pressure are not positions in the real world. It's a "state space" or "configuration space". At any moment in time, I could measure my system's temperature and pressure, and plot a point at (temperature(t), pressure(t)). As the system changes through time according to whatever rules govern its behaviour, I could take snapshots and plot those points (temperature(t+1), pressure(t+1)), (temperature(t+2), pressure(t+2)). This would give a curve "trajectory" that represents the systems evolution over time.
Okay, that's a 2D state space. But imagine I had a simulation of 10 particles (maybe some planetary simulation for a game). For each point I have maybe a 3D position (x,y,z) and a 3D velocity (vx, vy, vz). So I need 6 numbers to fully describe the state of each particle, and I have 10 particles. Therefore to fully describe the state of the whole system, I need 60 numbers. I therefore have a 60-dimensional state space. But each of these dimensions does not represent a position measurement along some axis in the world. In fact, only 30 of them do (3 * 10), the other 30 represent velocities.
For an abstract perspective, try Sheldon Axler's Linear Algebra Done Right.
For a more concrete perspective, Gilbert Strang's lectures: https://www.youtube.com/playlist?list=PL49CF3715CB9EF31D
In the context of neurons, while the neurons are in the 3 spatial dimensions, the connections of each neuron can be encoded in a feature vector. Each connection can specialize on one feature, e.g. the hair color of the person. These connection features can be encoded in a vector. The number of connections becomes the dimension of the vector. Not to be confused with the physical 3D spatial dimensions of the neurons.
The nice thing about encoding things in vectors is that you can use generic math to manipulate them. E.g. rotation mentioned in this article, orthogonality of vectors implies they have no overlap, or dot product of vectors measures how "similar" they are. Apparently this article shows that different versions of the sensory data encoded in neurons can be rotated just like vector rotation so that they are orthogonal and won't interfere with each other.
Linear algebra usually deals with 2 or 3 dimensions. Geometric algebra works better on higher dimension vectors.
y = a1 * x + a2 * x^2 + a3 * x^3 + a4 * x^4
where you only have one input and one output, but 4 constants that can be adjusted. These 4 constants make up a 4D vector.
Think about how a CNC machine works, you can have CNC with more than 3 axis, for example a 4 axis CNC machine can move left/right up/down backwards/forwards and also have another axis which can rotate in a given plane.
From a more mathematical perspective just think about the number of parameters in a system (excluding reduction) each parameter would be a dimension.
I have found it easiest to think of a logical dimensions or configurations when thinking of higher dimensions. Physically it can be a row of bulbs (lighted or not) wherein N bulbs (dimensions) can represent 2^n states in total. The 2 here can be increased by having bulbs that can light up in many colours.
Smartphones eg. measure six dimensions of freedome, including rotation about every axis. 3 for location, 3 for orientation.
this has very little to do with synapses.
The comment I am replying to, your comment in the tree, and the one next to you, does not seem to match that request in any sense.
Now, simplified definitions are an art, but Feynman managed it with Quantum Electrodynamics -- so it is not impossible to do it for complex subjects. And it seems to me the less you understand a subject, the less simple and more confusing your explanation will be, such as the explanations given by the other posters here. (fyi: I do not understand enough to properly convey my understanding clearly -- which is why I have not attempted to do so)
The "connections" you mention aren't the issue, in my understanding of the biology. Neurons are already very strongly interconnected by numerous synapses, so they already do physically fit together in their available 3D space, and appear capable of representing high-dimensional concepts. (See caveat below.)
The "higher dimensions" here are not where the neurons exist, only what they're capable of representing. If we think about a representation of the concept of a "dog" for example, there are many dimensions. Size, colour, breed, temperament, barking, growling, panting, etc etc. Those attributes are dimensions.
Take two dog attributes: size and breed. You can plot a graph of dogs, each dog being a mark on the graph of size vs breed. Add a third dimension and turn the graph into a cube: temperament. You can probably imagine plotting dogs inside this three dimensional space.
It's very difficult to imagine that graph extending into 4th, 5th or further dimensions. And yet, you can easily imagine, say, a dog that's a large, black, friendly Labrador with a deep bark who growls only rarely. We could say that dog can be represented as a point in 6-dimensional space (or perhaps a 6-dimensional slice through a space with even more dimensions, just a slice through 3D space could produce a 2D graph).
The number of connections between neurons may be related to the number of dimensions they can represent. In honesty, I don't know, and I guess that if there is a relationship it may not be linear. So neurons might be capable of representing 4 dimensions with fewer than 4 synapses, for example, I don't know. Seems possible to me, though.
Caveat: I think my reasoning here may be fallacious: "the fact that neurons are capable of representing high-dimension concepts demonstrates that they have adequate synapses to do so". It seems akin to anthropocentrism, I'm not sure. Perhaps it's just a circular argument. I think it provides an adequate basis for an ELI5 though.
I look forward to further comments!
In the visual cortex, neurons are arranged in layers of 2D sheets, so that perhaps gives an extra dimension to fit connections between layers.
If I remember correctly, the integers Z form spaces, too. Z^2 can be illustrated as grid, where every node is uniquely identified again coordinates or by two of its neighbours, eitherway v = (a, b).
Adjency lists or index matrices are common ways to encode graphs. My modelnof a neuron network is then a graph.
I imagine that, since Neurons have many more Synapses, that's how you get a manifold with many more coordinates.
Each Neuron stores action potential much like color of a pixel and its state evolves over time, but that's when the model becomes limited.
How it actually represents complex information in this structure I don't know.
PS: Or very simply put, physics has more than three dimensions.
You can multiplex in frequency and time. I'm not sure if neurons do it, but it's certainly possible with computer networks.
There is evidence for sparse coding and PCA-like mechanisms in the brain, e.g. in visual and olfactory cortex [2,3,4,5]
There is no evidence though for backprop or similar global error-correction as in DNN, instead biologically plausible mechanisms might operate via local updates as in [6,7] or similar to locality-sensitive hashing [8]
[0] Sparse Autoencoder https://web.stanford.edu/class/cs294a/sparseAutoencoder.pdf
[1] Eigenfaces https://en.wikipedia.org/wiki/Eigenface
[2] Sparse Coding http://www.scholarpedia.org/article/Sparse_coding
[3] Sparse coding with an overcomplete basis set: A strategy employed by V1?https://www.sciencedirect.com/science/article/pii/S004269899...
[4] Researchers discover the mathematical system used by the brain to organize visual objects https://medicalxpress.com/news/2020-06-mathematical-brain-vi...
[5] Vision And Brain https://www.amazon.com/Vision-Brain-Perceive-World-Press/dp/...
[6] Oja's rule https://en.wikipedia.org/wiki/Oja%27s_rule
[7] Linear Hebbian learning and PCA http://www.rctn.org/bruno/psc128/PCA-hebb.pdf
[8] A neural algorithm for a fundamental computing problem https://science.sciencemag.org/content/358/6364/793
(and when the neocortex that does most of the processing with this data is actually closer to a very thin, almost two-dimensional manifold wrapped around the sulci)
There has to be an information-theory connection between the physical form and the dimensionality of the memory lookup, even if they aren't referring to precisely the same thing, right?
The average neuron has 1000 synapses, and for geometric reasons (Synaptic connections take up space) most of those are to other neurons that aren't very far away in 3D space.
Similarly: yes, physics limits neuronal connectivity. The actual space of neuronal connections lies on a manifold inside the full “n squared, divided by 2” dimensions of connectivity of any old set of n points. That still doesn’t mean neurons can’t represent high-dimensional concepts, because your treatment of physical dimensions as the same thing as concept space is still mistaken. Taking your 1000 synapses number for granted, the input to a given neuron would be 1000-dimensional, not three. If you’re not arguing the concept space is 3d, and merely arguing against those who’d say neuronal connectivity isn’t limited by physical constraints, then I’d advise a reread of the ancestor comments; none of them are saying that.
Nature stumbled onto the path that it did because we don't have high enough nutrient food or fast enough neurons.
[0] http://www.scholarpedia.org/article/Grid_cells
[1] Time (and space) in the hippocampus https://pubmed.ncbi.nlm.nih.gov/28840180/
[2] Organizing conceptual knowledge in humans with a gridlike code: https://science.sciencemag.org/content/352/6292/1464
I'm not convinced the author's analogy of cross-writing to fit more information on a page is actually going to be helpful to most people's understanding. It led me at least to try to imagine visually what's going on, to picture the input being physically rotated. This is more akin to the more abstract but inclusive concept of rotation from linear algebra, where more dimensions (of information, not space or time) makes sense.
I thought that was more of a case of a human's facial recognition being a special function, and we're not able to process two or more people's faces at the same time. Like, see the details in them, recognize that it's their face.
You're either looking at one person, or the other, but if you try to look at both of them at the same time, they become "blurry", unrecognizable, even though you remember all the other information about them both.
But that's not related to memory integrity and new emotions/sensations?
Like with those chords in a research. Mice hear one chord, and by association from memory it expects other chord. But instead it hears some third chord. Expected and unexpected chords have perpendicular representation, if I understood correctly.
Here you see a picture, and expects one interpretation or other. You have memory of both, but you get just one.
Possibly it doesn't apply, I do not know. I'm trying to understand it. The obvious step is to make a prediction from a theory, should interpretations oscillate, if it has something to do with perpendicularity of representation in neurons?
When I hear another chord instead of a predicted one, do prediction and sensations oscillate? I'm not quick enough to judge based on a subjective experience.
I figured out how to change it at will eventually, if you close your eyes then open them and look at the bottom of the picture first it’s an old woman. Do the reverse and it’s a young woman. Eventually you can do that without the eye closing step but never would I say I could see both at once.
Just rapidly switch.
Very interesting!
This happens to me often.
Once I'd seen it once, the mother-in-law is now prominent. I can still see the wife if I concisely choose to, but the mother-in-law is now the default, strange huh?
I showed my wife the picture and she couldn't see either woman until I pointed out features. Interesting!
I only see the young woman before I became disinterested in making the other one happen because why
https://www.simonsfoundation.org/2021/04/07/geometrical-thin...
This sounds like the early conservation of momentum / conservation of energy debates. (Not that they used those words back then.)
I guess that would sort of be like the opposite of DRAM - cells maintain state when undisturbed, but the "refresh" operation is lossy.
https://www.npr.org/transcripts/788422090
Quote (although it’s missing context if the full show):
> Yeah, I think it's really interesting. I think it's really interesting to think about why we do these things, why we misrecollect our past, how those kinds of reconstruction errors occur. And I think about it in my own personal life - I share my memories with my partner. And many of us who have partners, we have these sort of collaborative ways in which we recollect. But those collaborations often result in my incorporating information into my memories that were suggested by this individual, but I never experienced. And so I might have this vivid recollection of something that only my partner experienced because we've shared that information so often. And so that's how we can distort memories in the laboratory. We can just get individuals to try and reconstruct events over and over and over again. And with each reconstructive process, they become more and more confident that that event has occurred.
Or like any analog data medium ever :)
here are some references
https://pubmed.ncbi.nlm.nih.gov/?term=memory+reconsolidation...
(AFAIK it's totally wrong, but I really like it anyway. I hope there is another specie in the universe that use it.)
Word embeddings frequently encode particular traits in different 'regions' of a 256(ish) dimensional space. AFAIK, It is also why we think of element wise addition (merging) in neural networks as an efficient and relatively loss-less computation. The aggregation after attention step used in Transformers (GPT-3) fundamentally relies on this being true.
Although from my reading, there is an inherent assumption of sparsity in such situations. So, is it reasonable to assume that human neurons are also relatively sparse in how information is stored ?
but I found the preprint of the paper on biorxiv.org: https://www.biorxiv.org/content/10.1101/641159v1.full
Hahahaha!
What are these numbers, precisely?
If you're asking about the exact numbers here's a snippet from the xlsx document. ``` ABCD_mean ABCD_se ABCD_mean ABCD_se XYCD_mean XYCD_se XYCD_mean XYCD_se day neuron subject time 0 6.012574653 0.5990308106 6.181361381 0.5737310366 6.59759636 0.6419092978 6.795648346 0.5716884524 1 2 M496 -50 ```
According to the article SEM neural activity, though this is way beyond my ability to interpret.
Assumes facts not in evidence? I feel it’s incredibly common for memories to intrude on perception of present, and be rewritten by new experiences.
The problem is suppose you have 4 neurons that need to understand a memory and present experience at the same time to make a decision (for example that you see a hot stove and memory that hot stoves hurt). The incoming neurons from memory and experience each have 4 neuron connections, which fire at some rate. Lets represent this as the firing rate of the neurons per second in a 4d vector:
experience: <1.0, 0, 0, 0> (1 pulse per sec on axis 0) memory: <1.0, 0, 0, 0> (1 pulse per sec on axis 0)
If you "add" these together at the downstream neurons, you won't be able to tell which was a memory and which was sensation. A simplified explanation of how neurons work is by combining voltages from their incoming neurons. Example:
downstream sees: <2.0, 0, 0, 0>
upstream could be a memory with <1.0, ...> and experience <1.0, ...>, or memory <2.0, ...> experience <0.0, ....>, or memory <0.5, ....> and experience <1.5, ....>. There are many possible vectors that could "add" to produce the downstream effect, so it makes it harder for those neurons to "learn" the pattern.
As a math equivalence, if I ask "what two numbers sum to 10", there are many solutions (its impossible to disentangle the original numbers).
To make it easier to learn these patterns, what if we used only separate elements of the incoming vectors to represent this information (so the elements of memory and experience could be seperated)?
So some intermediate neurons can transform the representation. We can constructor orthogonal vectors (since the vectors above are sparse):
experience: <1.0, 0, 0, 0> memory: <1.0, 0, 0, 0> => <0, 1.0, 0, 0> experience + memory: <1.0, 1.0, 0, 0>
The "memory" must undergo a "rotation" which moves data into an "unused" portion that won't conflict with the experience neuron firing pattern.
Now downstream neurons can use the data from each (its effectively merging memory and experience without confusing the signal). There is only a single memory and a single experience that combined will give the firing pattern, so the pattern can be learned.
Due to the way linear algebra works, its possible to do this with more complex numbers along arbitrary axes in an n-dimensional space (instead of doing it with a single axis/neuron and all others being zero).
For a physical corollary, imagine two images super-imposed on each other. If they are very distinct, you might be able to infer what the two source images were, but if they are similar it would be difficult. Now imagine a "lenticular" image that clearly displays two images by printing them at orthogonal angles on the medium. You can easily determine what content belongs to which image, but only having a single "print" to store the data (this isn't a perfect anology, but it illustrates the idea):
From the article:
> This use of orthogonal coding to separate and protect information in the brain has been seen before. For instance, when monkeys are preparing to move, neural activity in their motor cortex represents the potential movement but does so orthogonally to avoid interfering with signals driving actual commands to the muscles.
https://www.biorxiv.org/content/10.1101/2020.10.26.356089v1
"Our findings reveal that the specific spectral tunings of the four cone types near optimally rotate the encoding of natural daylight in a principal component analysis (PCA)-like manner to yield one primary achromatic axis, two colour-opponent axes as well as a secondary UV-achromatic axis for prey capture."