The problem is not only context or subtext, it is even worse from an entirely predictable standpoint:
Consider a large corpus of text (books, articles, ...).
1) Concepts that are WELL UNDERSTOOD BY HUMANS will not be explained by humans when they reference them: the word or concept "pet" (as a verb) will show up in many sentences refering to the petting of cats, dogs, horses,... and ML will correctly predict the conjunction of "pet" with any of these words in the sentence. It will even be able to train ML to confabulate realistic sentences in the sense of echolalia. Consider then the sentence:
"The lady pets the cat"
The computer will recognize that the presence of cat is not surprising (bingo!). The computer will have no idea that this probably involves one or more cycles of the lady's hand gently pressing down on the fur belonging to the cat, then while still preessing down, moving the hand in the 'natural direction of the hairs' (NOT the other way around) probably from closer to the head towards the tail, and probably lifting the hand before repeating the cycle so as to massage the cat or perhaps so as to remind the cat of its time as a kitten being licked by the mother cat.
No book or conversation in the corpus will give this detailed description exactly because humans expect each other to understand this.
2) Concepts that are POORLY UNDERSTOOD BY HUMANS will be vigorously (but often erroneously) explained by humans communicating to each other what they think is going on: endless texts about religion, sexuality, perpetuum mobiles, economy, ...
How do we even expect the computer to produce a sane result, even if it correctly guesses the context?
That said, I do believe relatively helpful natural language processors to be possible, but they will have to be vigorously trained by multiple human curators individually analyzing a sentence and trying to find (probably true) statements about what a sentence implies:
Starting again with:
"The lady pets the cat"
One curator might mention hands touching fur while moving.
Another curator might add that one can also conclude the hand probably presses down on the fur.
Yet another notes the sentence also implies the lady is still alive, for else she would not be able to pet.
The first curator now adds that the cat as well is probably alive, for else the lady would probably not want to pet the cat, since massaging is useless to a dead cat.
The third curator now mentions that the sentence implies one or more cycles of an individual pet stroke.
Etc...
As you can see this quickly becomes an expensive operation.
Now one might train an adversarial neural network to look at a sentence (or sentence with context) and a list of probable conclusions, to predict if the list of valid conclusions is complete or incomplete. And then only send the incomplete ones to humans?