The deepest problem with deep learning
medium.com
medium.com
Symbolic AI fell out of favor primarily because it was not delivering results in impactful problem areas. Deep learning is currently popular because we are nowhere near the limit of what results it can produce.
Can this change? Of course! The history of deep learning itself proves as much. But if you want to genuinely influence the direction of the field, you have to lead by example and produce novel/interesting research results, not by kvetching in The New Yorker that your favorite approach is not getting enough attention.
That's why the article discusses examples where currently popular approaches fail.
> not on their conformance to a (vague, incorrect, untested) model of human cognition.
What model of human cognition is "incorrect"? And is it the one presented by Marcus, or a strawman?
> But if you want to genuinely influence the direction of the field, you have to lead by example and produce novel/interesting research results
Are you claiming Marcus has produced no interesting research results?
> not by kvetching in The New Yorker that your favorite approach is not getting enough attention
Why use the term "kvetching"? I'm curious.
Not the OP, but I'm not familiar with Marcus' contributions. What would you consider his top contributions? Are they all purely theoretical, or is there something that's already been applied in the real world?
The model of human cognition I'm referring to is the hybrid connectionist-symbolic one that Marcus is well known for advocating (are YOU strawmanning? lol). I'm criticizing it for being more a theoretical model than one grounded in the physical realities of the brain, which of course no one really understands. Proposing a research program on that basis requires a high burden of proof.
> Are you claiming Marcus has produced no interesting research results?
Yes I am claiming that, if the benchmark for "interesting" is deep learning.
There are indeed areas where deep learning is limited, and hybrid approaches could be superior. I would argue that there is not even close to enough evidence that a hybrid approach has improved generalizable power.
> Why use the term "kvetching"? I'm curious.
Huh? I guess it's the term my mother would use.
No-one in machine learning, yeah, because machine learners mostly don't take neuroscience classes ;-).
> I'm criticizing it for being more a theoretical model than one grounded in the physical realities of the brain, which of course no one really understands
You're contradicting yourself. On the one hand, you claim Marcus' model is "incorrect". On the other, you claim there's insufficient evidence either way. Which is it?
> are YOU strawmanning? lol
Do you know what the term "strawmanning" means? What could I possibly be strawmanning since I was asking for clarification?
> Proposing a research program on that basis requires a high burden of proof.
As opposed to...?
> Yes I am claiming that, if the benchmark for "interesting" is deep learning.
"Deep learning" isn't the correct benchmark since that's what Marcus is critiquing (to some extent) in the first place.
> I would argue that there is not even close to enough evidence that a hybrid approach has improved generalizable power.
Then you'd be wrong. Here's a good place to start your research:
http://science.sciencemag.org/content/331/6022/1279
> Huh? I guess it's the term my mother would use.
Your mother taught you to describe scientific debate as "kvetching"? That's disappointing.
In particular, your comments have broken this guideline: "Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith."
Can you clarify? How have I "been doing it a lot"? And how was my comment "in the flamewar style"?
> Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize.
Which part of the GP's comment do you think I'm unfairly interpreting?
HN threads are supposed to be thoughtful conversation—not cross examinations or verbal boxing matches, let alone setups to make other people look bad.
I see absolutely nothing wrong with the second comment you linked to, in particular. What do you find objectionable about it?
I'm also still curious to know which part of the GP's comment you think I'm strawmanning.
When I read your comments, they seem to carry a charge of aggression presenting itself as inquiry. In order to be in the spirit of this site, that quality needs to go. Here's an old pg quote that I love, which expresses the right spirit: Comments should be written in the spirit of colleagues cooperating in good faith to figure out the truth about something, not politicians trying to ridicule and misrepresent the other side.
Re facts: your comments are probably being flagged, first because they do a lot more than merely state facts (examples: "I encourage you to open your mind", "Can you not read?", "The only one being a jerk here is you"), and second because facts can also be weapons; it depends on how they're used.
Re the GP: taking someone to task for using the word 'kvetch' is cheap and has overtones. You can see that in the bewildered way the other user responded. Had you been following the site guidelines, you'd have dropped that bit. Please take the guidelines to heart, and you'll do much better here.
> The problem with all these questions is that the cost of answering them is an order of magnitude greater (more, actually) than the cost of posing them.
No, I don’t think so. Requesting clarification and/or evidence when someone makes a claim like that is not unreasonable. It’s
(1) ensuring everyone’s on the same page to avoid pointless misunderstandings, and
(2) upholding the burden of proof, i.e. extraordinary claims require extraordinary evidence.
Genuine discussion would be impossible without these two things. If someone makes an extraordinary claim, they should have to spend time doing the necessary research and presenting it.
> I encourage you to open your mind.
I see nothing uncivil about this.
> Can you not read?
> The only one being a jerk here is you.
Neither of these appear in the comments you linked to.
> and second because facts can also be weapons; it depends on how they're used.
Who gets to decide whether “facts are used as weapons”, and what does that even mean? This is veering into dangerous territory and I’m surprised to hear you think this way.
> Re the GP: taking someone to task for using the word 'kvetch' is cheap and has overtones.
I did mean to take them to task for using the word “kvetching”. Because debate/critiquing/disagreement is not “kvetching”, and it’s absolutely unfair for people to characterize it as such. Debate is essential to science. There’s nothing “cheap” about that.
Of course facts can be used as weapons. When one middle schooler points out the physical awkwardness of another, the more factual it is, the more cruel it is. "I'm only stating facts" is an evasion—why those facts and not others? There are infinitely many facts. The ones to mention don't select themselves; we do that, and there's a lot more going on in us than just "facts".
Fair, I didn't notice that. Calm down.
> Of course facts can be used as weapons. When one middle schooler points out the physical awkwardness of another, the more factual it is, the more cruel it is. "I'm only stating facts" is an evasion—why those facts and not others?
Because they were relevant to the discussion?
> There are infinitely many facts. The ones to mention don't select themselves; we do that, and there's a lot more going on in us than just "facts".
Which of the facts I stated do you think I shouldn't have stated?
My understanding is instead that symbolic AI was working pretty damn well for its time. Expert systems routinely outperformed experts, for sure. The AI winter that killed them was brought on by political decisions taken by people who didn't really understand the field.
Here's a good read on that historical period:
Avoiding another AI winter, editorial in IEEE Intelligent Systems.
https://www.computer.org/csdl/mags/ex/2008/02/mex2008020002....
Btw, "Intelligent Systems" is such a funny little expression. Basically, it was used by AI researchers during the 80's AI winter to be able to get funding for their work; because they wouldn't get any if they called it what it was, AI.
Both failures ultimately were caused by not enough computing power. Even though Deep Learning and Convolutional NNs look like major advances today, they never could have been practical before about 2005: There just wasn't enough computing power.
If modern computer power were thrown at symbolic AI the same way it's been thrown at NNs, it highly likely symbolic AI would experience similarly-impressive gains.
What's the basis for this conjecture? Is there a mathematical model for symbolic manipulation that would benefit from parallel execution/GPUs the way ML applications do?
I don't know whether rule search or logic unification could be mapped onto GPU or TPU operations; I suspect not, but it's worth looking into.
For instance (also in another comment) Decision Tree learners basically learn a set of If-Then-Else rules. They're one type of symbolic machine learning and there's more where they came from (e.g. Ross Quinlan's FOIL, for First-Order Inductive Learner, which is basically a first-order version of decision trees; Inductive Logic Programming which I study for my PhD; and many, many more). This work has dwindled, but it's still going.
So, no, you don't have to write rules by hand, anymore than you need to set the weights of a neural net by hand. You can just learn them.
https://en.wikipedia.org/wiki/Combinatorial_explosion
The vulnerability of early logic-based AI to combinatorial explosion was the main argument against funding AI research put forward in the Lighthill Report, the document that shut down AI research in the UK in the 1970s and contributed to the AI winter on the other side of the Atlantic, also.
Obviously, today we have more powerful computers so combinatorial explosion is less of an issue, or anyway it's possible to go a bit further and do a bit more than it was in the '70s.
One area of symbolic AI that actually does benefit from parallel architectures (though not GPUs) is logic programming with Prolog. Prolog's execution model is basically a depth-first search, which lends itself naturally to parallelisation (one branch per search). Even more so given that data in Prolog is immutable (no mutable state, no concurrency headaches).
But, in general, anything people did 20 or 30 years ago with comptuers can be done better today. Not just symbolic AI or neural networks. I mean, even office work like printing a document is faster today and that doesn't even depend on GPUs and parallel processors.
Oops, sorry. Meant "one branch per processor".
I really hate these kind of arguments where some people have strong beliefs that "it would work" and not do it or invest themselves. Its like being in a pub and listening people talk about politics while neverhave had any responsibilities
a good area for examples is human games (e.g. chess, Go, Atari games, etc.) Symbolic AI has been pushed hard but has lost definitively to deep learning. Furthermore symbolic approaches had decades of investment, compared with less than a decade for the deep learning approaches.
Another good area for examples is natural language. Marcus admits that deep learning is the only viable approach to "speech understanding" (which really means transcription). He doesn't mention translation which demands a lot more "understanding" and where deep learning excels relative to symbolic approaches, again with decades of investment on the symbolic side and much less on the deep learning side.
Except for inherently symbolic problems like theorem proving, I can't think of any AI domain where deep learning doesn't dominate or seem likely to dominate symbolic approaches.
First Marcus responds to Bengio's interview with the arrogant and pretentious "I told you so" tweet, instead of simply saying that he agreed with Bengio.
Then LeCun tweets a snarky and disrespectful response, instead of simply saying that Bengio's comments are different than Marcus's earlier critique.
As a result of these tweets, I lost a huge amount of respect for both scientists.
"I stand by that — which as far as I know (and I could be wrong) is the first place where anybody said that deep learning per se wouldn’t be a panacea, and would instead need to work in a larger context to solve a certain class of problems. Bengio was pretty much saying the same thing."
which stakes out a priority claim (that deep learning is not a panacea) that appears somewhat superficial given the huge literature on causal modeling and shortcomings of the alternatives.
So, LeCun's response isn't about Marcus != Bengio, it's about the fact the Marcus' critique fundamentally hasn't deserved response or recognition this whole time, because there's not a constructive way to engage. That's a totally fair point to make, though a guy in LeCun's position probably should have said, "Gary, none of us have ever thought Deep Learning was (already) an answer for everything. The primary work of this community is to make it better."
> And it’s where we should all be looking: gradient descent plus symbols, not gradient descent alone.
That's the most specifically defined proposal that I see in the article. It is a reasonable mission statement to fire up your lab or community about. It's also not something outside the ongoing deep learning discourse, and it's a direction I personally am excited about. There has been great work recently by DeepMind [1], for example, about using gradient descent to theorize about symbolic relationships.
However, the statement above is also nowhere near operationalized enough to say "agree / disagree" in any scientific sense. A specific model demonstrating advantages of the marrying symbolic and gradient-based reasoning (c.f. the one above) would open itself to productive discussion of its successes and failures. People have asked for this (including in the tweetstorm), but most weirdly Marcus seems to be responding that operationalizing a model or even a success criterion is itself a waste of time! Quoting:
> I actually think benchmarks are to some degree the wrong way forward, and have said that for two decades. People need to take a step back, and reflect on where things stand, rather than rushing into the next bakeoff.
I couldn't disagree with this statement more, and it brings to my biggest personal axe to grind about the "is this how we get to AGI" query: AGI itself hasn't been defined in a way that is compatible with empirical discourse. Machine Learning at a term in many ways exists to abandon "AI"'s association with AGI, because the community realized that focusing on success at specific, objectively measurable tasks would move them forward more effectively.
When people focused on AGI say that ML in its current form won't get there, my answer is: "It's almost tautological that we're going to need new methods to get to something called AGI, but we can't even discuss progress on it until you give a measurable definition".
I don't blame folks (including Marcus) for wanting to continue discussing AGI as an abstraction - maybe they will be the ones to find a good operational definition! But it's weird and unscientific to say people shouldn't continue work on other directions in the mean time, or to say that a specific technique is essential before either the technique or the end goal has been effectively defined.
[1] https://deepmind.com/blog/neural-approach-relational-reasoni...
On second thought, I wish I myself was always very thoughtful and respectful in responding to questionable statements :)
Since we have DL but don't yet have AGI, it's quite obvious to everyone that we are missing something.
I repeat: the interesting question is precisely what.
I repeat my claim: DL (defined as gradient-based learning of non-linear functions) will be part of the solution.
I tend to think of deep learning as adding additional sensory capabilities to machines. In this sense, is a network misclassifying a school bus as a snowplow significantly different from a GPS sensor generating garbage position data due to multipath? All existing sensors are noisy and can get confused, and yet no one says that GPS isn't useful because it can be so completely wrong in urban canyons. Of course, no one says GPS is going to become sentient - this brings us back to my first question above.
Ng wasn't even really saying it in a way that is explicitly making the case - at least not enough to cause this kind of fracas.
It's a bit unfortunate really, and I'm not sure what's driving the animosity or aggressiveness around this issue. While I agree with Marcus, I don't think anyone actually disagrees - largely because nobody is arguing what he is trying to address. Marcus has built a strawman and is tilting at Windmills a bit.
However I think it's a reasonable argument to make that a General Intelligence may use some form of a symbolic system. I'd even accept the claim that General Intelligence would require some form of expert systems to cross the threshold. So in that sense, combining DL with Expert Systems, other symbolic methods or approaches like Bayesian Causal analysis is probably required.
One of the biggest missing pieces in my mind has always been, what drives it? What does this theoretical AI want/need, and how does it know to maximize on those values?
No, this has been a straw-man argument, made repeatedly in popular press for public recognition, for years. Even the NYT article to which Marcus responded in 2012 [1] does not make the claims to which it responds! Here is the sentence in Marcus' New Yorker post where it transitions to attacking a claim nobody made (hint - look at where the quote ends):
> While the Times reports that “advances in an artificial intelligence technology that can recognize patterns offer the possibility of machines that perform human activities like seeing, listening and thinking,” deep learning takes us, at best, only a small step toward the creation of truly intelligent machines.
As a rule, the academic deep learning community avoids discussions and claims about what constitutes a "truly intelligent machine", instead focusing on results in specific applications which can be objectively measured.
I'm sure you can find both academics and public press that claim deep learning is a panacea. But it's disingenuous punching up to continually imply that leaders in the field are slow to acknowledge its limits.
[1] https://www.nytimes.com/2012/11/24/science/scientists-see-ad...
A trend in the DL/ML research community is the strong and vocal distancing from discussing or approaching the problem of General Intelligence or anything which might look like sci-fi AI. You can see this with aggressive railing against silly robots like Sophia and others, which have gotten press.
I think this is warranted, and understandable as the field is so scared of another AI winter.
On the other hand it leaves the field without a real progress vector. Especially given that broadly, ML research isn't scientific as such. In other words, the vast majority of "research" is not around hypothesis testing and using computing to test questions about intelligence. Rather it's demonstrating increments of improvement within very narrow computing tasks.
I make that point not to diminish the work or results, as they've obviously been astounding. Rather, to point out that unlike basic or social sciences research it's not attempting to surface fundamental truths (theories, laws etc...) about some natural or emergent phenomena, but rather continuously benchmarking on narrow tasks - often without prior reference. Now that ILSVRC is gone, I'm wondering what kind of benchmarking is going to emerge as the new vector for the field.
My personal time projects are mostly combining symbolic AI and deep learning, but I am still trying to find a non-Python solution, including Haskell TensorFlow bindings, Armed Bear Common Lisp with DL4J, and exporting trained Keras models to a Racket environment - all plausible hacking environments but none feel ‘just right.’ If you are working on the same ideas please get in touch with me.
EDIT: make that a double thanks, Flux looks great: concise and expressive.
I like TensorFlow with Keras much more than using DeepLearning4j, but DL4J is callable from Armed Bear Common Lisp. I have simple examples in my book that you can read free online: https://leanpub.com/lovinglisp Unfortunately the example in the book is a little weak because I wrote the deep learning code in Java and just called my Java code from Common Lisp. I should modify the example to build models, train, and evaluate using the DL4J APIs directly from Lisp.
EDIT: in the book example, I just call train and evaluate entry points in my Java code from Lisp so I can train new models from Lisp but can’t create new models with new architectures without modifying the Java code. I really should change that.
Java itself tends to be fairly verbose I know, but not only are there wrappers but for inference if you want to use python for basics and then us for inference. We enable several ways to do that. Folks tend to oversensationlize these "us vs them" framework "fights".
Rather than do that, why not just use what's good for each use case? Python for training, java for deployment?
Beyond that, you could also consider building your own wrapper maybe? In clojure, there's jutsu.ai for example: https://github.com/hswick/jutsu.ai
I use Keras at work so that is why I prefer it - I am very used to Keras, and to a lesser degree TensorFlow.
Just tonight I was experimenting with jutsu.ai, and your advice of training on Keras and running on DL4J is good advice. I hacked up a little code to import Keras weights into Racket, but I can only process dense layers, and not RNN or complex architectures.
EDIT: thank you so much, great talk! Even though the Juji system does not directly link TensorFlow models into Clojure, the architecture of using Kafka to manage low level black box model calls (they treat trained models as functions as I do in my Racket hacks) is a good idea. I also like the rule/template language they designed. Great work.
To cut a very long story short, machine learning took off in part as an attempt to automate knowledge acquisition for expert systems, the dominant AI paradigm at the time. Somehow, for reasons that are not entirely clear to me, the goal of learning the rules for a rule-based system was abandoned and since then research has focused almost exlusively on just learning.
As a result of this, most machine learning work today does not consider what one can do with a trained model. And yet, just having, say, an object classifier, is not very useful on its own. For example, a robot car must be able to take decisions based on the objects its machine vision algorithms identify. Although there are learning techniques that consider this particular problem, i.e. navigation, there is little work on general inference and reasoning over the output of trained models.
And this is not a symbolic vs statistical thing. I work in symbolic machine learning, and we are pretty rubbish at that, too. It's like we have lots of little pieces of a puzzle all lying around and nobody is really trying to put them together. Instead, we each carve our little (or bibger) niche and pretend nothing exists outside of it.
Why? It's a mystery to me. Gary Marcus is doing good to try and shake up the field a bit. We need to move on from our successes just as we move on from our failures. And he's damn right about the how, too: symbolic AI is due for a comeback.
____________________
[1] For instance, you can see that in the foundational theoretical text for modern machine learning, Leslie Valiant's A Theory of the Learnable where the learning framework is described in terms of the propositional calculus, with a concept represented as a predicate that recognises vectors of boolean variables as members of itself or not.
And of course, probably the most famous representative of propositional logic learning is the class of algorithms we probably all know as decision trees; they learn disjunctions of conjunctions - propositional logic rules.
Just another note here to say that the original artificial neuron, the Pitts and McCulloch neuron, was designed as a logic circuit with the ability to learn its own function. At the time -the 1930's- the propositional calculus was considered to be a good model of human intelligence.
How would we define "symbolic", though? I'm about to give a presentation on some related matters next week, and I've been trying to think about how to avoid just saying, as Marcus is perceived to do, "lol psychology critiques AI so AI needs to use symbols."
For example, there are loads and loads of symbolic AI techniques which don't demonstrate basic features of "symbolic" human cognition such as causal inference and productivity/compositionality, on top of symbolic techniques having essentially no way to address the Frame Problem.
The root cause of said Twitter firestorm was Marcus’s accusation that world’s talent and money is moving en mass in deep learning research which is likely to turn out dead end for achieving “real” AI. Even worse, many uninformed decision makers think AI is already a done deal due to all the media hype.
The counter point to his accusations were that deep learning is what works now and pretty well for many real world problems so we shouldn’t be putting it down. It is likely that deep learning might become integral part of broader some Artificial General Intelligence framework in future. So no harm in keep improving it. This would be especially be important when most of the accusers don’t actually have realistic alternative with much to claim.
In short, there is a real risk that researech funding will go to support research that doesn't know its elbow from its knee, done by people who don't understand what they are doing (because they don't understand what was done in the past). That is not a rosy situation.
I get the counter point- I'm myself a bit annoyed by Marcus' insistence on how deep learning is not true AI. Very few people really expect to see "true AI" in our lifetimes. A classifier that works well is probably the best we can do and it's very reasonable that there is such excitement about the fact that we can do it. But if we spend too much effort developing classifiers, well, we'll have great classifiers. And nothing else.
And then- what?
https://twitter.com/tdietterich/status/948811917593780225
https://cdn-images-1.medium.com/max/800/1*W5WhToR_WP4I74Lo4u...
Lecun didn't say he is not allowed criticize. He said that his contributions are zero but his criticisms abound.
Also some of the arguments are odd: It is not possible to show the limits of deep learning because people don't know how to prove what those limits are. If there was a "limits of the universal approximation theorem" paper^, then people would use that to derive the limits of their DL systems. OTOH he proposes a number of possible implementations for symbols for which the evidence is not there, and there is no theoretical reason that necessitates their existence. Both arguments are not really falsifiable.
^ Actually there is at least one, however it seems it requires knowledge about the smoothness of the approximated function. https://www.sciencedirect.com/science/article/pii/S0888613X0...
Historically there seem to be three views. One is the classic symbolic manipulation view, the second is what seems to be the ml view, and then there is the hybrid view that Marcus advocates.
The first two views seem to have in common believing that their single approach alone would be sufficient for AGI. Marcus thus differs from both of them in that he thinks neither is sufficient.
Assuming Marcus is taking the right approach, the question then becomes whether the hybrid approach could reach AGI. In my opinion it could not, though it probably would be able to out-think human beings in some important ways.
Frozen images at that. I’d expect better results from systems trained on videos of moving objects.
If deep learning is mimicking how our brains work, wouldn't it be theoretically as powerful as our brains? In terms of architecture and complexity of the model.
Neurons are families of cells that come in many shapes and most of the time get oversimplified for didactic purposes. The processes involved in neuroscience are also oversimplified most of the time.
It is often said that dendrites are the input, axons the output, and that spikes are fired when a threshold is met... but that's not always true. The biological situation is full of special cases.
Then you have the way the network is formed. The amount of neurons and layers in a deep neural networks are just hardcoded in advance. That's not the case in biology.
Then you have volume: biological neural networks have many more neurons, and many more connections between those. For example each Purkinje cell (in the cerebellum) can have over 200k dentritic connections.
We aren’t mimicking how our brains work in any broad sense. At best it’s a very narrow definition.
In this case - the larger the matrices - the more you can do. There’ll probably be a Moore’s law of AI at some point.
Why would you assume that? I did a cursory GIS and Bing search, and somewhere between 1-5% of the returned images were bunless wieners.
Literally the first image returned for hot dogs is a bunch of bunless ones on a grill. There are plenty more.
For comparison, I tried out Clarifai's image classification API: https://clarifai.com/demo It has no problem classifying some bunless ones on a grill (https://www.hot-dog.org/sites/default/files/2016-09/sausage%...) as being sausage/pork/beef/hotdog, with nary a 'sandwich' to be seen, while conversely, a Big Mac (https://d1nqx6es26drid.cloudfront.net/app/uploads/2015/04/04...) gets sesame/burger/lettuce/sandwich/bun/bread tags with no sausage/bratwurst/hotdog to be seen. Clarifai's NNs seem to do fine on 'hot dog or not hot dog'...
the idea originally was that a neuron can be roughly modeled as something that sums up inputs from other neurons and then outputs some kind of threshold/activation of those inputs. this is probably sort of true of biological neurons although as living cells they are much more complicated than that.
but lots of recent and highly successful techniques in deep learning have given up on even that very weak amount of biological imitation. One simple example: lots of techniques involve multiplying together the outputs of "neurons". There is no sensible biological interpretation of this operation.
Neural networks are Turing-complete (assuming all the usual pedantic assumptions you need to make that statement meaningful) so in that sense they are as powerful as our brains.
The other way is more controversial, but I don't think it's outrageous to speculate that a brain could be accurately simulated by a Turing machine.
Fair enough. I got a different impression from reading your original comment :)
That's of course true, but then most computers aren't that similar to a classic Turing machine either. The idea of Turing machines is that they can simulate anything else that we mean when we say "computation".
> Can a Turing machine simulate a brain - as in, create intelligence like the one our brains expose?
Yes, it can. I mean, if it can't, that means there's some notion of "computability" that a Turing machine doesn't capture - this would violate the "Church-Turing" thesis. Such a thing is possible, of course, since we have no proof that a Turing machine is the limit of what we mean by "computable" - but it would be a completely revolutionary fact in Computer Science, and one that is very unlikely, considering how many years people have been trying to come up with alternative models of "computation" that all turn out to be equivalent to Turing machines.
-
The human brain isn't magic - it works according to the laws of physics like anything else. Either it is simulatable by a Turing machine, which would make the most sense, or it isn't, in which case we're missing something about how you can compute things, which would be surprising and amazing.
Take for example machine vision. ConvNET is inspired by the lowest levels of human visual cortex and there are analogies to be drawn, but no direct mimicry.
https://neurdiness.wordpress.com/2018/05/17/deep-convolution...
how about no gradient descent at all