The problem is we don't understand how the brain is architected at anything other than a very coarse level. That architecture is the product of millions of years of evolution so we have some catching up to do.
And more generally, it’s pretty clear that there is a mixture of context-sensitive inhibitory and excitatory signals both deriving from the dendrites (inputs) and from the chemical stew of hormones and so forth they find themselves immersed in.
There are neurobiologists who claim (controversially, I concede) that a single neurone is a fully self-contained ‘processor’[1], and not the equivalent of single logic gate.
Those who somehow confuse our “neural network architectures” with the actual biology in our skulls are really misguided. Based on the evidence, one can at best opine that a TPU-based matrix add/multiply unit is approximating a very clunky and un-nuanced neural network. To assert that they are one and the same is... indefensible, when one goes beyond even the coarsest detail. A neural network as we implement it technologically is nothing but a weighted network of threshold units. That is all.
[1] To clarify: some kind of highly context-aware DSP with both long- and short-term storage facilities that allow it to react not only to the inputs that it is being fed, but also to inputs it was fed previously (and possibly outputs it produced previously as well).
To be fair, many lectures/books/articles/blog posts on ANNs start off by presenting a biological neuron, and saying that ANNs approximate this.
E.g. https://youtu.be/uXt8qF2Zzfo?t=385 (MIT 6.034 Artificial Intelligence, Fall 2010. 12a: Neural Nets)
I started off along much the same path (I bought and read Neural Networks: A Comprehensive Foundation (1994) by Simon Haykin) when I was in high school, read it all and thought I knew everything there was to know on the topic, and then started to read about neurobiology and was stunned to discover I was living a lie. Then I read The Computational Beauty of Nature (1998) by Gary William Flake and decided I was going to dedicate myself to understanding the complexities we brush under the rugs.
I think the Hacker News/Silicon Valley/Machine Learning/Developer crowd really need to look beyond the virtuosismo of their implementations. It’s a very introverted and self-referential mindset.
You might like this article: https://blog.piekniewski.info/2018/08/28/fun-numbers-about-t...
Then, gradient descent is a centralized and supervised learning algorithm that has nothing to do with all the decentralized and emergent processes that take place in biological brains. For example, consider all the complex interactions of the various neurotransmitters and neuroreceptors.
Biological brains (let alone human) are actual physical realizations of networks with very complex and specific topologies. New connections grow and are formed in a very real sense, and this process interacts with the environment in a variety of ways, plus it is spatially embedded.