It all boils down to this paragraph, and no one can say it is not true.
It all boils down to this paragraph, and no one can say it is not true.
Hofstadter, having completely missed the boat on connectionism, committed himself to a failed paradigm, and been wrong about AI time after time (remember when Hofstadter claimed beating humans at chess would require computers to have emotions and be able to write poems?), totally fails to grapple with what NNs do do. For starters, Google Translate is a free service and not necessarily SOTA, even when he wrote this. For example, if you were going to ask 'what do multi-lingual translation seq2seq RNNs (or Transformers) understand?', the immediate obvious thing to discuss would be questions about the 'interlingua' which their embedding yields, things that word embeddings learn (including surprising things like relative locations of cities or sizes of objects), or about transfer learning to tasks that require grammar or reasoning or common-sense like Winograd schemas like in GLUE or SuperGLUE, and how they improve massively over earlier approaches and how they are scaling (one of the most important trends in AI). There is a great deal which could be said here, demonstrating why it works so well and why it doesn't work sometimes, and an AI researcher is just the sort of person to ask.
And unfortunately, Hofstadter does none of this and instead bangs on about some errors he found, like it's a total blackbox and understanding is a binary where you either like the same poetry Hofstadter does or you are a cretin, and takes the attitude that any error shows that the approach has failed, is failing, and always will fail because it is 'deeply lacking' (what, exactly? And to think, he complains about 'deep' being abused for rhetorical purposes before and then goes and produces a shallow exercise like this).
I thought this was a lousy neo-luddite essay in 2018, and it's only gotten worse since then. In another 10 years, it'll look like Cliff Stoll.
From what I see and understand, even if you have multiple layers of perceptrons and back propagation. It is still really "deeply lacking" because you are not able to train networks with all context there is. You can still overtrain network, so more training is making it actually worse. You can train NN only so much, just like you cannot teach dog to speak, it can perfectly fine understand your commands but you will never communicate with the dog on a level of "this painting is beautiful, you agree?".
I also do not believe in possibility of creating AGI.
I don't know what you're trying to say here. NNs can have plenty of context. Let's take just the Transformer architecture for concreteness, ignoring all the others. Something like GPT-2 uses a fixed window, but it can be expanded arbitrarily widely if you're willing to pay the compute price; the scaling can be improved within that window by not computing attention over the full window, per Sparse Transformers; recurrency can be added to make it more RNN-like, as does TransformerXL; compressed versions of the memory can be stored and fed back in, allowing lookback arbitrarily far, per Compressive Transformers; multi-modal input, like images + embeddings + text, are of course completely possible and can represent arbitrary metadata and context; there are many different ways to add in external memory stores (DeepMind regularly experiments with these, like the DNC); and so on. What exactly do you think is impossible in principle in the connectionist paradigm to make NNs do (and why do you think your brain is exempt from this)?
Maybe some parts of image recognition require spinal cord neurons and muscle tissue to interact? So is it just count of inputs to neuron, or there are ways where inputs to neuron are modified by other chemicals in a body? Are we making weights on inputs stronger because of good reasons, are making some other inputs weaker because of good reasons?
I study a class of machine learning algorithms that learn logic programs (as in Prolog, or ASP) from examples. The field is Inductive Logic Programming. ILP algorithms are characterised by their ability to incorporate background knowledge and impose strong inductive biases on the hypothesis space to learn accurately from a handful of examples.
My favourite example of this is learning a grammar of the aⁿbⁿ language. My implementation of one of the algorithms that my research group studies (I'm a Phd research student) is Thelma:
https://github.com/stassa/thelma
In Thelma's README on github you will find an example of Thelma learning a grammar of the aⁿbⁿ language from three positive examples and no negative examples. The learned grammar generalises to any n. In typical CFG notation it's this grammar:
S → AB
S → AS₁
S₁ → B
A → a
B → b
(The nonterminal S₁ is invented, i.e. it was not given in the original
learning problem defintion. The preterminals A and B are background
knowledge.)By contrast, neural networks can only learn fragments of this grammar up to some limited n and that only if they're given tens of thousands of training examples.
For instance, a classic paper by Gers and Schmidhüber [1] claims LSTM "generalisation" from training sets of 22 to 42 thousand aⁿbⁿ strings, but their average generalisaion is, e.g. from n in [1,50] (i.e. that's the value of n in training strings) to n in [1,430] (n in testing strings) and their "best" generalisation is at most to n in [1,1000].
On Thelma's README on github I have a small testing query that tests how the aⁿbⁿ grammar it learned from 3 positive examples with n in [1,3] generalises to an aⁿbⁿ string of e.g. n = 100,000:
?- _N = 100_000, findall(a, between(1,_N,_), _As), findall(b, between(1,_N,_),_Bs), append(_As,_Bs,_AsBs), anbn:'S'(_AsBs,[]).
true .
Well, it's a correct grammar so it generalises perfectly. You can bind _N to
any number your computer memory will allow and it will still parse.So, besides the plug of my work (sorry) I would say that the earlier paradigm that you say has "failed" with respect to recent connectionist success is alive and well and it still gets many things right that the connectionist paradigm struggles with mightily, in particular robust generalisation from very few examples (without expensive pre-training etc), and of course any task that requires reasoning [2].
______________
[1] "LSTM recurrent networks learn simple context free and context sensitive languages": https://ieeexplore.ieee.org/document/963769 See table 2 on page 7 for the results I quote above.
[2] The examples of good performance on reasoning tasks you bring up have been strongly criticised, e.g. in https://arxiv.org/abs/1907.07355 (Probing Neural Network Comprehension of Natural Language Arguments) or Gary Marcus' recent paper on a critical appraisal of deep learning etc.
P.S. I'm very sorry to see you're being downvoted.
I'd say that it's rather statistical machine learning research that is in a bubble, satisfied to solve the same old classification problems again and again, apparently oblivious to the fact that there are many problems that their algorithms can't touch because they are neither differentiable, nor reducible to classification.
Unfortunately, since statistical machine learning is the current dominant paradigm, success is measured in terms of what it can do- and anything it can't do is either completely ignored or discounted.
Or, like I say: sour grapes.
I don't see why the distinction you are making matters.
> Neural nets can't solve it.
They can, it takes additional training to do so from scratch. Also, the work you cited is also nearly two decades old, I have no idea how SOTA would perform.
> solve the same old classification problems again and again
I truly have no idea what you're talking about, please point me to where, for instance, [0] has been solved "again and again"?
The claim I am making is not that Deep Learning can solve all problems, it cannot, the claim is that logic-based AI has failed to come up with anything of value whatsoever, whereas statistical ML has had sweeping results across many real problems -- for instance: AlphaZero solving Go, Chess, and Shogi; GPT-2 creating readable stories given a prompt; cars that can drive themselves under a wide range of circumstances. If you think a^nb^n is even remotely comparable to those feats, I don't think we can find any common ground.
On the Practical Computational Power of Finite Precision RNNs for Language Recognition
https://arxiv.org/abs/1805.04908
Typically, it claims generalisation but fails to demonstrate it.
As we agreed, it's a simple problem. Neural nets can't solve it. Do you understand why?
Re: classification, AlphaZero is a great example of what I'm talking about. The game-playing logic is provided by MCTS, a GOFAI algorithm (it's a variant of good, old minimax). The neural net component is only used to identify promising board positions, i.e. for classification.
As to chess, it was already solved by DeepBlue using alpha-beta minimax and an opening book of moves- another example of GOFAI. Even AlphaGo, AlphaZero's predecessor used domain knowledge in the form of example games played by humans.
But I'm getting the feeling that, while you are making this extraordinary claim that "logic-based AI has failed to come up with anything of value whatsoever", you are not very familiar with the history of AI research.
As a quick reminder, logic-based AI was the dominant paradigm in research for some 60 years. The big success of course were expert systems. As an early example, look for information on MYCIN, the first system to beat human experts at medical diagnosis (of infections, in partiuclar).
Even today, after the recent meteoric rise of deep learning, logic-based algorithms are state-of-the-art for many applications: in particular, planning, constraint solving, game playing (as in minimax variants) and anything that requires reasoning. In applied work, decision-tree learners, a class of machine learning algorithms learning propositional-logic models, are still widely used in data science work, more so than brittle and expensive to train deep neural nets.
Ugh. This sounds like a personal attack. I apologise unrservedly. I really did not mean to insult you or disparage your knowledge of AI research. There's much better ways to say what I wanted to say here.
Apologies, again. I can't edit my comment now.
I'm going with dooglius here. You are presenting a simple regular grammar as a success story of the GOFAI paradigm (which doesn't have that much to do with Hofstadter's paradigm, anyway), while the DL approach has been producing real-world successes on the hardest of natural language tasks. Your example is "not even wrong" in terms of criticisms like Marcus's, because it is so many orders of magnitude away from being useful or reaching the point where it even could fail or be criticized on those terms.
Despite myself, I am still surprised by the unwillingness of connectionists to admit this simple fact: neural nets generalise very poorly and have absolutely awful sample efficiency.
As to more complex, real-world tasks, this is a recent example on a language task using the same algorithm (but a different implementation):
Bias reformulation for one-shot function induction
https://dspace.mit.edu/handle/1721.1/102524
There's more of that in the literature.
It was my mistake to agree that aⁿbⁿ is a "simple" grammar. It looks simple and it's simple for a human.
But, aⁿbⁿ is a context-free language and as such it can only be learned "in the limit", i.e. by a number of positive and negative examples approaching infinity. It cannot be learned by positive examples only, not even in the limit. This is part of Gold's result [1], a famous result from Inductive Inference, the field that essentially predated Computational Learning Theory. Gold's result was that anything more complex than a finite language can only be learned in the limit and anything above regular languages requires both positive and negative examples.
Gold's result had a profound effect on machine leargning [2] and was a direct cause of the paradigm shift that came with Leslie Valiant's paper that introduced PAC Learning [3]. Valiant's paper placed machine learning on a new theoretical foundation where it's acceptable to learn approximately, under some assumptions (particularly, distributional consistency between examples and true theory). This is how machine learning works today and this is why machine learning results are accepted in scholarly articles, not because they are useful in the industry or make nice articles in the popular press.
Where aⁿbⁿ comes into all this is that it's one of a group of formal languages that are routinely used to assess machine learning algorithms. Keeping in mind that grammars for CFLs are impossible to learn precisely from finite examples, the point of the exercise is to show that some learning algorithm can learn a good approximation that generalises well from a small number of examples.
This is how formal languages are used in neural networks research also. For a recent example from the literature see [4].
The example in my comment shows instead that our algorithm can learn aⁿbⁿ precisely, not approximately, from only three positive examples. This should be surprising, to say the least. Gold's result says that this should not be possible, at all, ever, until hell freezes over. Why it is possible in the first place is another long discussion that I was hoping to have the chance to make here, but instead I allowed myself to be drawn into an adversarial exchange I should have avoided.
To put an end to it: neural nets have produced some spectacular results, and I don't want to discount them, but it is important to understand what those results mean, in the grand scheme of things. The fact that there is a machine learning paradigm that outperforms neural nets on a very hard problem (simple as it may look) means that the same paradigm may outperform neural nets on those other things that they do so well, like image recognition and speech recognition. Or it may just mean that different learning paradigms have different strengths that must somehow be combined (a more popular view for sure).
Anyway, sorry to spread noise and confusion in this thread when my goal was to do the opposite. I have a rubbish way of communicating.
_____________________
[1] Language identification in the limit:
https://www.rand.org/pubs/research_memoranda/RM4136.html
[2] And a few other fields besides. Noam Chomsky used Gold's result to argue for a Universal Grammar enabling humans to learn natural languages. This all is really groundbreaking stuff and a real shame that it's not more widely known in machine learning circles.
[3] A theory of the learnable:
https://web.mit.edu/6.435/www/Valiant84.pdf
[4] LSTM networks can learn dynamic counting:
I'm sorry but I don't understand this question. I'm not trying to do anything "for you". It sounds as if you think I'm trying to sell you something. I'm a researcher, I don't sell stuff.
But I'm curious: what are you currently doing with deep learning?
I'm a researcher too - I look for ways to build faster hardware for deep learning. But to answer your question - using deep learning I can classify objects, recognize speech, and translate language, as well as generate plausible text, nice sounding music, or beautiful pictures. Those are just few examples of what I can do with DL much better than with any other methods.
I find this unnecessarily provocative and it's certainly not what I said.