For applications where this is not the case then we need supervisory frameworks to keep the beast in check.
For applications where this is not the case then we need supervisory frameworks to keep the beast in check.
As I understand programs like BERT, they largely look at the words adjacent to a given word, without paying attention to any structure. Whereas in real human languages--any human language--there is a hierarchical structure. This has been known since 1957 (Chomsky's Syntactic Structures, although I suspect that context free or weakly context sensitive models are sufficient). Is there a way to force a neural net to construct hierarchical models? And of course the hierarchies are not different for every word--all nouns behave more or less the same, all intransitive verbs behave in a different way, etc. So the task of the neural net (or whatever learning system) is to discover the hierarchy for a given language.
Likewise, morphology is mostly suffixes and prefixes, with a handful of other possible structures (up to reduplication). And the affixes fit into a paradigmatic structure. The neural net should be primed to look for that kind of structure, and to back off to things like ablaut only if the affixal model doesn't work.
Is this a way that symbolic and neural approaches could play together? Let the symbolic channel and constrain the neural net.
Problem is that it does not scale to the sizes we need for dnn's. At least not yet.
http://www.optimization-online.org/DB_FILE/2020/07/7883.pdf
The idea is that you can express the network as an integer program. Now having that you have two big capabilities:
1) Write explicit constraints that the nodes need to satisfy. Aka if you activate this node, this and that must be activated too but that needs to be inactive. These constraints will be guaranteed to be satisfied when you get your solution.
2) You can solve the network configuration to proven global optimality or at least have a bound of how far you are from the optimal solution. That is in contrast to the current approaches with stochastic gradient descends that find locally optimal solutions with no information about how far you are from the globally optimal solution.
That being said, these problems are combinatorial ones, so scaling them up to millions of nodes would be challenging.
That's completely opposite to what happens. In BERT all words are related to all words so information circulates in one step between all pairs. This combinatorial interaction happens on multiple parallel "heads" and sequential "layers", so it can express complex symbol manipulation tasks.
In GPT information circulates between all pairs of words but only from past to the future, not the other way around (it is autoregressive).
What you are talking about is the RNN or LSTM. They only see the adjacent tokens. This causes an informational bottleneck which limits their expressive power.
> Is there a way to force a neural net to construct hierarchical models? And of course the hierarchies are not different for every word all nouns behave more or less the same, all intransitive verbs behave in a different way, etc.
Here are a few types of attention maps. When you stack them up you net a hierarchical model.