> I am genuinely curious why none of the recent Deep NLP approaches work here.
To my understanding, it's because deep learning is fundamentally aimed at pattern-recognition, but makes no attempt to have a model of "reality" underlying that recognition.
I've seen some people argue that these models of reality are, themselves, nothing more than higher orders of pattern-recognition, and therefore ought to be amenable to deep-learning type methodologies. I have my doubts. Deep learning requires reinforcement training from big datasets, and the process of modelling reality does not often accommodate.
The process of learning to see things, navigate through space, etc., is something that you do through reinforcement training -- more or less like a neural network. Newborn babies don't see "things": they see colours and lights and motion, and it takes a couple of years of training before they reliably classify those sensory inputs. That's very much a big data exercise: every instant your eyes are open, you're collecting more data and training on it.
In contrast, consider the schema "The council denied the protesters a permit because they [feared/advocated] violence." The reason you can parse that sentence is not because you have trained yourself on a dataset of thousands of councils and tens of thousands of protests, and can therefore now recognise which might be fearing violence and which might be advocating it. Instead, you've built a model of the world out of vastly sparser data, via structured logical inferences. Which just isn't remotely the way that deep learning works.
So current AI techniques seem to have gotten very good at matching (and surpassing) a subset of cognition, but I'm not convinced that they will scale to cover these kind of schemas. I do think that fundamentally new ML approaches will be needed for these kind of domains. Neural Networks seem to be good at cerebellum and occipital lobe sort of tasks, but we'll need something else for the frontal lobe.