I understand your points, but what I wanted to highlight was something slightly different: verbal communication and dialogs should still consists of magnitude more human conversation than text, or presentation, or basically anything digital.
Do you mean things like tone, facial expression, the general “energy” around a conversation, etc?
Basically, pick a random human from the street and they would be talking and listening a lot more than they would be writing. Those talking and listening is never going to be captured by any digital system.
It might end up that written text has enough similarity with verbal communication that it doesn't matter anyway. But that's hard to guess as a priori that it would be the case.