The conditions for causal inference being possible are pretty clear and have to do with the intentional modification of the local environment.
The intention to achieve some new environmental state, and your action to bring it about, is a dynamical activity that enables "deep" model building.
Causal inference is not going to be some "module" of the brain... it requires a body. When you place your hand on a hot surface, once, you immediately understand that it is hot. It does not require "induction" (as hume supposed). That is because our body identifies causes.
Its therefore pretty trivial to observe no NLP system understands language, or even can understand language, because it lacks this capacity to acquire language semantics via participation in environmental exploration. It has no body.
ie., you need to have experienced "on top", "green", etc. to know what "green leaves grow on top of trees" means. There is no meaning in the frequency co-incidence of symbols in text.
So no matter how much you are able to reproduce these patterns, they contain no content. The content is in the reader.