I study a class of machine learning algorithms that learn logic programs (as in Prolog, or ASP) from examples. The field is Inductive Logic Programming. ILP algorithms are characterised by their ability to incorporate background knowledge and impose strong inductive biases on the hypothesis space to learn accurately from a handful of examples.
My favourite example of this is learning a grammar of the aⁿbⁿ language. My implementation of one of the algorithms that my research group studies (I'm a Phd research student) is Thelma:
https://github.com/stassa/thelma
In Thelma's README on github you will find an example of Thelma learning a grammar of the aⁿbⁿ language from three positive examples and no negative examples. The learned grammar generalises to any n. In typical CFG notation it's this grammar:
S → AB
S → AS₁
S₁ → B
A → a
B → b
(The nonterminal S₁ is invented, i.e. it was not given in the original
learning problem defintion. The preterminals A and B are background
knowledge.)By contrast, neural networks can only learn fragments of this grammar up to some limited n and that only if they're given tens of thousands of training examples.
For instance, a classic paper by Gers and Schmidhüber [1] claims LSTM "generalisation" from training sets of 22 to 42 thousand aⁿbⁿ strings, but their average generalisaion is, e.g. from n in [1,50] (i.e. that's the value of n in training strings) to n in [1,430] (n in testing strings) and their "best" generalisation is at most to n in [1,1000].
On Thelma's README on github I have a small testing query that tests how the aⁿbⁿ grammar it learned from 3 positive examples with n in [1,3] generalises to an aⁿbⁿ string of e.g. n = 100,000:
?- _N = 100_000, findall(a, between(1,_N,_), _As), findall(b, between(1,_N,_),_Bs), append(_As,_Bs,_AsBs), anbn:'S'(_AsBs,[]).
true .
Well, it's a correct grammar so it generalises perfectly. You can bind _N to
any number your computer memory will allow and it will still parse.So, besides the plug of my work (sorry) I would say that the earlier paradigm that you say has "failed" with respect to recent connectionist success is alive and well and it still gets many things right that the connectionist paradigm struggles with mightily, in particular robust generalisation from very few examples (without expensive pre-training etc), and of course any task that requires reasoning [2].
______________
[1] "LSTM recurrent networks learn simple context free and context sensitive languages": https://ieeexplore.ieee.org/document/963769 See table 2 on page 7 for the results I quote above.
[2] The examples of good performance on reasoning tasks you bring up have been strongly criticised, e.g. in https://arxiv.org/abs/1907.07355 (Probing Neural Network Comprehension of Natural Language Arguments) or Gary Marcus' recent paper on a critical appraisal of deep learning etc.
P.S. I'm very sorry to see you're being downvoted.