I think the video stream in teleconferencing has absolutely zero redeeming qualities. People think it's there to be able to convey facial expressions and facilitate non-verbal communication, but I think it's a complete failure at that task.
For one thing, few listeners are actually looking at the speaker as the speaker is talking. They're most likely looking at themselves or some other thing going on with their computer, so their facial reactions are not really based on what the speaker is saying.
On the flip side, the speaker never gets to see people looking at them. Almost nobody looks at their camera instead of something on the screen, so the speaker never gets that "eye contact" feeling. Best-case scenario, you get a group of people trained to move the speaker video feed directly under the camera lens and they are diligent about making sure they are looking at the speaker. Even then, there is still a "20 yard stare" look to everyone. It also causes exhaustion as it puts you into a feeling that you're in an interrogation of some kind.
Additionally, it's such a narrow field of view for the camera. Non-verbal communication is more than just facial expressions, it's also body posture and standing distance. There are facial ticks that are also lost in the low-quality of the webcam feed, and the non-uniformity of every user's personal lighting settings creates an unnatural scenario where every person is lit differently than you'd expect, or from each other.
And finally, while teleconferencing has a lot of trouble with latency between when a person speaks and when the other people hear them, there is also a lot of latency between when you hear a person speak and when you see them. The audio and video feeds are not synced correctly.
By completely eliminating the video feed, conversations actually work a lot better. I get so many people who demand to have that video feed for the reasons they've been indoctrinated on, with little to no effort to even try audio-only conferencing.
And frankly, as a listener, I don't want the speaker to see my reflex reactions. I don't want them to see into my room. I really only want them to see the what I choose to let them see.
Thus, the avatars and the emoji reactions.