Study urges caution when comparing neural networks to the brain
news.mit.edu
news.mit.edu
This carries on into an extremely nuanced and technical discussion of the architecture of specific models used.
It's basically pointing out that in some previous papers, authors thought that grid cells always arose when solving this problem, but in fact this only occurs when specific implementation choices are made. So those papers were incomplete, and the phenomenon isn't as clearcut as before.
However! This new complexity still tells us something; if only certain architectural choices produce grid cells, then brains must (in some sense) implement those architectural choices. And the models that don't produce grid cells must be doing something differently to how the brain does things.
In summary I think this paper is probably saying a lot less than most people here are reading into it; some papers are accidentally oversimplifying, and we've found more complexity that needs to be explained. More thorough hyperparameter-space exploration can identify brittle results. It's not some deep point about whether it is philosophically or logically consistent to compare deep NNs to the brain.
Interlacing isn't 4-way or 6-way, it's 10e3-way, and each interlaced connection has a weight that's nonlinearly time-dependant based on how long since last firing.
Every cyclic connection is potentially a self-sustaining oscillator.
None of these features are efficiently implemented in current silicon.
"Caution when comparing neural networks to brains" is underselling it. They're profoundly different kinds of network; nobody is building (or even publicly planning) silicon that has that kind of interconnect breadth.
The reasons that there aren't much more fully connected layers, is that this doesn't work. Actually, one of the key developments in NNs is architectural, minpools, ReLus, U-net. All are key for modern networks, all architectural.
It is, of course, very common to perform (eg) convolution with a larger kernel size, or to use a dense layer.
However, unlike wet-neurons:
* Convolution has the same local shape for each cell. * Convolution has no self-suppression for recent activation, vs time-dependent, nonlinear response in wet cells. * Current silicon offers no performance advantage for interconnect to adjacent cells (could be done).
With ~80 billion neurons in a brain, 1000 is not like a dense layer at all.
What I understand is that he claims the underlying algorithms that govern our behavior and how it evolves from birth are ingrained in our genetic code. Current neural network models try to model our behavior, but it is way behind when it comes to discovering those ingrained algorithms.
Human brains can be tricked too, but never this way and never beyond our capacities for rational thought.
Humans can be tricked very easily by optical illusions and it isn't uncommon for some illusions to be intentionally built to 'harm' people (eg patterns on the floor which make you lose your sense of balance). Even with rational thought such things can be difficult to deal with. We're probably just as vulnerable to adversarial attacks, the issue being that unlike artificial neural networks we don't have an easy feedback loop to run millions of times in guiding a similar adversarial search.
Not even dynamic ones.
Magicians will tell you they're fooling you, but con artists can use many of the same patterns and you only find out too late.
Like getting him to spend $44 billion for Twitter?
(An ex of mine is convinced that Musk is a con artist, but she's also a literal card carrying anarcho-communist; I'm not that cynical about Musk).
Even at a lower level, I had my bank[0] call up and tell me there was too much money in my account and they'd really recommend a wealth management consultation to avoid me being scammed, and that wasn't even £100k.
That said, I was thinking mainly of street cons — shell games, possibly even shoplifting and pickpocketing — as the previous discussion was about optical illusions. Business level scams are about a broader category of cognitive bias, and I'd say almost all gambling is that type of thing, likewise bitcoin, dulce et decorum est, and populist politics.
[0] or at least they said they were, but I said no before getting to the point where asking for proof the call wasn't itself a scam would've been useful
It's inconceivable to me that humans wouldn't be trickable by exactly the same sort of adversarial inputs-- it's just that because we're not differentiable there is no feasible way to find these inputs.
People have constructed fairly impressive optical illusions based on our understanding of the neural structure of the early stages of vision processing. The fact that we lack more complicated examples like "random images" that make us feel hate or disgust or that we're convinced are our mother is simply due to our lack of understanding and access to the higher neural structures.
And no, this standard practice does not eliminate adversaries.
I think probably that this kind of meta-understanding exists on a continuum where fruit flies and slime molds exist on one side and current AI exists somewhere in the lower third and we exist somewhere in the upper third and a future AI with a vast number of AI brain components and huge amounts of training will eventually exist at the other end, possibly as far from us as we are from the fruit fly.
We do not understand our data.
In general yes, I believe most people will only accept thinking machine when it can reproduce all our pitfalls. Because if we see something and the computer doesn't, then it clearly still needs to be improved, even if it's an optical illusion.
But our bugs aren't sacred and special. They just passed Darwin's QA some thousands years ago.
I'd agree about sacred, but I have a hunch they may indeed be special… or at least useful. Current AI requires far more examples than we do to learn from, and I suspect all our biases are how evolution managed to do that.
Humans get a lot of data.
AI gets more examples.
Tesla autopilot has a movie of every second it's active, for every car in the fleet that uses it. It has how many lifetimes of driving data now? And yet, it's… merely ok, nothing special, even when compared to all humans including those oblivious of the fact they shouldn't be behind a wheel.
And yet we're living in the age of misinformation where propaganda spreads like never before. Isn't that essentially an adversarial attack which shows how susceptible humans are to them too?
For the record, i don't think that's NN are brains either. I think our issue to differentiate sufficiently between them is because we don't really know what sentience really is.
I simply don't understand why literally everyone immediately jumps at CNNs, RNNs (transformers et al.) -- they're extremely expensive, slow and definitely not usable for SIGINT-sized intel projects.
[1] https://arxiv.org/pdf/2112.14820.pdf [2] https://arxiv.org/pdf/1402.2902.pdf [3] https://arxiv.org/pdf/1509.08255.pdf [4] https://arxiv.org/pdf/2205.15407.pdf [5] https://arxiv.org/pdf/1602.05925.pdf [6] https://arxiv.org/pdf/1511.08855.pdf
Because for text they work a lot better.
HTMs are competitive with other non-NN techniques on small datasets (see [1] you listed) but nothing particularly amazing.
I'd speculate this is because they use bag-of-word variants like LSI and TF-IDF. Prior to Transformers this was a competitive technique and you could get state-of-the-art results on most things using similar techniques.
This representation of data matters a lot. Even just switching to a word embedding representation and using a SVM or something gives a decent gain in most circumstances.
But transformers are much better, particularly on harder tasks (eg question answering on long documents). You can't really see how significant this difference is on these small datasets, but as an example BertGCN is getting over 89% accuracy (HTM in [1] gets 83%).
It's possible (likely!) some of this gain is from the better representation Transformers use, not just the model.
> definitely not usable for SIGINT-sized intel projects
If SIGINT in this context means signal intelligence (on text data) then I assure you that they are being used.
[1] https://arxiv.org/pdf/2112.14820.pdf
[2] https://paperswithcode.com/paper/bertgcn-transductive-text-c...
> If SIGINT in this context means signal intelligence (on text data) then I assure you that they are being used.
Maybe, I don’t know what projects you’ve been involved with. For the terabit-level pre-sorting of SIGINT data they’re absolutely definitely not used. If at all on the selected information of interest. My information concerns intel actors in Europe.
Often these systems will have really bizzare artificats, people with 3 arms, etc. However at the same time when you glance at the output without looking carefully you will sometimes miss these artifacts even though they should be absolutely glaring.
I suspect also the reason the images look OK at a glance is because the images as a whole also represent patterns in the model so they actually come from "real life" / artist created images and thus have some sense of cohesion. But making the AI have all the right patterns so it never makes a mistake at all scales of the image while also being able to combine the pattern with real understanding of what they are conceptually is the real trick but until then it will be a "salad bowl collage" thing at random intervals.
The closest thing to the brain it looks like to me is simply the hierarchical nature of it which seems similar to v1/v2/the vision system in humans but I've only been told that, I'm no neuroscientist.
FWIW I don't think there is anything particularly wrong in the model architectures or training data that in some fundamental way makes it impossible to always get 2 arms. After all, lots of other tricky things are almost always correct. I suspect it's a question of training time and model size mostly (not trivial of course as it's still expensive to re-train to check modified architectures etc). It's also a matter of diffusion sampling iterations and choice of sampler at inference time, for the case of SD.
I also don't think there's anything wrong with the model architectures in themselves or the data, nor that it is impossible, only that it is hard and as you say I think it needs a lot of data and clever engineering to fix mistakes. It may even be possible to fix most mistakes, over time, which would be pretty impressive imo, but the absolute limits of what a model can produce/"contain" with our hardware is kind of an open question though interesting.
But it rarely would put out say 8 arms. And the repeat artifacts are miles ahead of earlier stuff like clip draw or disco diffusion. So it does seem to have some idea of what's going on, just isn't perfect yet. It gets much worse without the 512x512 resolution, if you push both dimensions it loses scene coherence a lot more.
However, where it struggles I find is with finer details, and also _placement_ of things like arms, eyes, and relationships between them. This I think is because it only has a general idea of the shape of persons but no data for the exact specifics like where the arms, legs, eyes and so on should be placed in a very realistic anatomical way, and this is where I think the challenge is - the gap between a general pattern of a person and an extremely specific but also general one where it can modify it and transform it like a real human artist can. I'm not sure that's in the data exactly
I’m no NN guy, but to me all it seems as basically underconstrained and unrelated to “understanding”. It’s like these e.g. woodwork, magic trick, dancing, guitar, etc teachers who fail to message a way to do something and can only tell “look”, then just do it, ask you to repeat, and get annoyed when you fail again.
This is a fundamental misunderstanding of what it is doing.
You can see in work like https://twitter.com/lintool/status/1579830653126086656 that the model does have an understanding of what parts of the visual model represent as concepts.
And you're right that this is pretty unfounded intuition. Humans often seek meaning in things without meaning, so it might be unfounded. At some point all i can really do is shrug and say it feels "spooky" to me.
That is thoroughly confused to the point of uselessness.
The reason you get structural issues is because it's hard for the architecture to express large scale structure, but they get better and better at it simply by scaling up the network.
But now you get SD, dalle and others which add more information not just by scaling, but also by mapping sentences/words to pre-existing images that already have cohesion. That way when you write in sentences to the text prompt, the model has more semantic information about what an eye is, but (IMO) only _indirectly_ because it will map a sentence to images that match that sentence. The question is always what information is actually contained in the training set and what is missing from it and when it creates an image where is the information from etc.
In some ways, that means I think that meaning to us as humans, is different from scaling which is almost like pixel resolution except resolution of patterns and differentiation of patterns. Meaning in this sense is things like creating a doorway with no actual door, but still the doorway itself looks super realistic is rendered. You can fix it by scaling and increasing the differentiation of patterns I guess, but you can never fix all instances completely with scaling. That's why in some ways I think meaning is sort of orthogonal to scale, however on a philosophical level, they should converge but that's for another topic.
I may have missed something in my thoughts here because this is sort of difficult to talk about without writing a book eventually.
There's simply nowhere for the computation to go.
Turns out nobody quite knows how to draw a bicycle. They get the gist but the details don't make sense.
I strongly suspect that if we do ever fully map the "architecture" of the brain, the result will be a massive graph that's not readily understandable by humans directly. This is already the case in biology. We'll end up with a computational artifact that'll help us understand cause and effect in the brain, but it'll be nothing like a tidy diagram of tensor operations like in state of the art ML papers.
There’s some image I see on occasion that’s 100% garbage. If you focus on it you cannot make out a single thing. But if you glance at it or see it scaled down, it looks like a table full of stuff.
I don't know if AGI is down the road diffusion models have taken us. I'm not even really sure what most people mean by AI when they talk about it. But stable diffusion et al are clearly super human. I'm not sure that AGI is down the trail cut by diffusion models, but if it's ever accomplished, these models will almost assuredly represwbt some of the learnings required to get there.
Very well then I contradict myself,
(I am large, I contain multitudes.)
They keep telling me this and, yet, I can't stop doing it. The more I learn about neural networks, the more I feel like I understand my own brain (whether accurate or not). And conversely, the more I think about thinking, the better my theories about how I'd build ML-based system to solve specific problems (admittedly, most untested). Neural networks seem like too useful of a model to simply give up because they aren't completely accurate.
Of course, this is all just for personal use - mostly introspection. I wouldn't exactly do medical work based on the model.
Neurons have a lot going on, they send and receive signals through a multitude of mediums, not just neural impulses, and they're capable of plasticity when it comes to the connections they make between other neurons. Neurons also don't have simplistic activation functions, they're capable of doing a lot more with the information they receive and send. Also, gradient descent and back propogation don't take place in any part of the brain.
Through that lens, I see NNs as if they're like really complex and impressive Markov chain generators. They can produce results that look intelligent, but it's just statistical correlations, and not at all how the brain works.
Without defining what's essential, I'm nervous to call the comparison insufficient. If a topological subset of neurons isn't good enough, what do we need in addition/instead? If we stuff NNs full of complicated (how complicated?) activation functions, does that new system do the trick? Or add...47 new "neuron" variants? Or swap the learning scheme from gradient descent to something fancier? (For that matter, do we even know what the brain's scheme is, and why GA/back prop isn't an acceptably extremely crude approximation of it?)
The brain is so unimaginably intricate. Our models are hilariously simple in contrast, of course. But what of those mismatches are differences in kind vs. differences in magnitude?
"All models are wrong. Some are useful"
-George E. P. BoxWhat if this a more accurate representation:
"It's like the connection between dictionaries in real life and dictionaries in programming."
The neuron could be implementing NN theory in a way that is optimized for it's environment.
Funnily enough, that’s how I see brains. They start out as a few neurones basically implementing ‘hard wired’ logic, then some feedback loops form and next thing you know they’re asking “why am I here?”
These "just so" stories are attractive but it is quite important to realize a metaphor which is intuitive and you perceive as useful is nothing at all like the process for finding real scientific truth. There is also a lot of introspective value to modeling the world as being controlled by mysterious gods who are pleased or appalled at your behavior and that's why good and bad things happen. Perhaps useful for some people but nothing at all like truth.
When I said "I can't stop," I was referring more to this tendency to borrow models to explain unrelated systems. It's just a thing my brain wants to do and I can't help it (and again, I seem to convince myself that it's somehow accurate or useful even if, rationally, I'm quite sure it's not).
The usefulness of neural networks has not ceased, despite researchers' early ideas and hopes about its biological analogies having somewhat sheared away.
The paper also does not have any reference to a study or paper that explicitly states that a neural network is a good model for grid cells. (Please correct me if I am wrong.) So I am left wondering why this direction was chosen.
Maybe it's a little cynical, but this topic seems to have been chosen (at least in part) to produce a splashy headline. Or in other words, to give the Stanford and MIT PR engine something to print.
This is the sort of obvious thing we all knew to be true. Why people with access to lab animals and a fully stocked microbiology lab needed to prove it (again) I do not understand.
It seems it is at least possible, that there is speed-of-light quantum communication within the brain. And that consciousness may hinge fundamentally on this. If this is true, we're pretty much back to square one in terms of understanding.
[1] https://science.ucalgary.ca/news/state-consciousness-may-inv...
Chemistry is quantum physics at its core. It is just that quantum equations are so hard to solve for anything bigger than hydrogen that most of the times, chemists prefer to use empirical rules to do their job.
What does this even mean?
That's why Einstein thought it was spooky! But in the widespread interpretation[1] of quantum entanglement it turns out not to be a problem because (while entanglement effects are real) it's impossible to transmit information or action via it.
Worth noting that this link doesn't talk about that at all. Instead it's about quantum chemistry effects.
[1] https://en.wikipedia.org/wiki/Copenhagen_interpretation#Acce...
I really don’t get why everyone wants the Brian to operate on some new QM effect other than peoples perception that a 100 year old theory is somehow cutting edge, spooky, or something. Perhaps it’s that the overwhelming majority of people who talk about QM don’t actually understand it even a little bit. Odd bits of QM are already why lasers, LED’s, and transistors work. You use incites from the theory everyday in most electronic devices, but it’s just as relevant for explaining old incandescent bulbs we just had other theories that seemed to explain them.
When you talk about QM a a theory of how the world operates, there are wide ranges of QM. Everything from predicting the structure and energy states of a molecule, to how P/N junctions work, to quantum computers. Now, for the first one (molecules), the vast majority of QM is just giving ways to compute the electron density and internuclear distances using some fairly straightforward and noncontroversial approaches.
For the other ones (P/N junctions, QC computers, etc), those involve exploiting very specific and surprising aspects of quantum theory: one of quantum tunnelling, quantum coherence, or quantum entanglement (ordered from least counterintuitive to most). We have some evidence already that there are some biological processes that exploit tunnelling and coherence, but none that demonstrate entanglement.
Personally, I think most people think the alternative to Penrose- the brain does not compute non-computable functions, and does not exploit or need to exploit any quantum phenomena (expect perhaps tunnelling) to achieve its goals.
Now, if we were to have hard evidence supporting the idea that brains use entanglement to solve problems: well, that would be pretty amazing and would upend large parts of modern biology adn technology research.
Your other points are based on such fundamental misunderstanding that it’s hard to respond. Saying something isn’t the output of classical computing processes while undemonstrated, is then used to justify saying they must therefore use Quantum Phenomenon. But logically not everything that is either classical or Quantum so even that logical inference is unjustified. Logically it’s like saying well it’s not a soda so it must be a rock.
PS: If people where observed to solve problems that can’t be solved by classical computer processing that would be a really big deal. As in show up on Nightly News, and win people Nobel prizes big. Needless to say it hasn’t happened.
I should have said "problems which do not have computable solutions" rather than "set of problems computable by a quantum computer", which seems fairly pedestrian compared to what Penrose is saying.
As to the specifics, let’s just say there’s a reason he was publishing books rather than peer reviewed papers.
The functional channels for neurons are well understood, even if we're still diagramming out all the types of neurons. Voltage gated calcium channels are pretty damn simple in the grand scheme of things, and they don't leave space for quantum interactions beyond that of standard molecular interactions.
The only part of the brain we don't understand is how all the intricacies work together, because that's a lot more opaque.
Well, okay. But so far there's no evidence to suggest the brain uses quantum effects discovered/verified fairly recently.
The "common" properties of electricity and chemistry where fairly well known and modelled well by 1900.
Therefore it's obvious that they're related. So it seems clear that AGI will be solved with the help of quantum physics.
My aunt Mildred is a very well renown academic and has written much on this topic. She unfortunately is also not well understood. So it seems quite clear - perhaps obvious - that AGI will be solved by applying some Mildred.
A study urges caution comparing Jellyfish to Jelley ... tasters found they are not the same (even though I hear that fried jellyfish taste nice...)
study urges caution comparing the model to the real thing, as the model has some generalizations the real thing does not ...
my assumption, the author hides the rather technical contribution of the paper behind a tautology to get some attention. seemed to have worked on hackernews as it's on the front page.
The brain is a hard drive but the body is the whole computer.
Science is proving physical causation. Not just writing down what we want to be true.
I like his idea of finding 0-days in physics. :)
> Unique to Neuroscience, deep learning models can be used not only as a tool but interpreted as models of the brain. The central claims of recent deep learning-based models of brain circuits are that they make novel predictions about neural phenomena or shed light on the fundamental functions being optimized... Using large-scale hyperparameter sweeps and theory-driven experimentation, we demonstrate that the results of such models may be more strongly driven by particular, non-fundamental, and post-hoc implementation choices than fundamental truths about neural circuits or the loss function(s) they might optimize. Finally, we discuss why these models cannot be expected to produce accurate models of the brain without the addition of substantial amounts of inductive bias, an informal No Free Lunch result for Neuroscience. In conclusion, caution and consideration, together with biological knowledge, are warranted in building and interpreting deep learning models in Neuroscience.
And IMO a succinct description of the problematic assumption being cautioned against in the study's introduction section:
> Broadly, the essential claims of DL-based models of the brain are that 1) Because the models are trained on a specific optimization problem, if the resulting representations match what has been observed in the brain, then they reveal the optimization problem of the brain, or 2) That these models, when trained on sensibly motivated optimization problems, should make novel predictions about the brain’s representations and emergent behavior.
---
I think to most, the problem with claim number 2 directly above is obvious, but it's important to also look at claim 1.
Who cares? Everyone knows ML models do not reflect the mechanics of how biological brains work at low level. The most obvious is that they use electricity, discrete numbers, much faster refresh rate etc. As a consequence the other low level "implementation details" will differ. The closer to "the hardware" the more differences there will be. I woukd be extremely surprised to see similar encoding, activation waves/patterns as in biological systems in ML for this reason, but also because how different the learning data and even the learning mechanism is. The brain has no backpropagation.
However, there is deep similarity between both and IMO we are not far from AGI(decades at most). There is a measure of similarity between some advanced ML models (stable diffusion in visual, bloom in reasoning) and how our thinking works. This is especially visible when those things break or produce unexpected results in comparison with damaged/psychedelic human brain.
Just like a human performing a math calculation and a computer performing the same calculation are doing essentially the same thing despite vastly different "implementation method", and same as computers helped us advance our understanding of mathematics(and physics etc) ML models will help us understand more about how our own thinking works.
Just as there is something universal in an act of adding two numbers, there is something universal in an act of processing language to derive intent and carry out complex instructions.
The crucial unknown however at this stage is whether our most advanced ML models are indeed using the same universal high level mechanisms we do to understand our input when they demonstrate their incredible capabilities or are they simply some advance method of compressing and searching through the training data? The first stage of answering this question is to determine if there is really a difference. Perhaps all we are, are databases doing an effective search algorithm over our training data?
This is what science hopefully will answer in coming years. In one way the pace of incredible discoveries of those new and bigger models is not leaving the scientific community enough time to study those models fuly. I can imagine many lifetimes could be spent just studying bloom or stable diffusion, but how to do it when new models twice their size show up 6 months later? How to focus on one model and one application of it in this quickly changing environment?Still, I'm very grateful that I can see this progress during my lifetime. While growing up in the 90s I had this feeling of "missed opportunity" that I never saw nor I have taken any part in the computing revolution that happened before I was born, but this new AI revolution certainly makes up for that.