Objects mostly just sit there, or move in more predictable ways. But when it comes to faces and bodies, we're super-tuned to facial detail, and are heavily primed to read emotions and context from facial details and from posture.
So if any part of that falls into uncanny valley, the whole experience looks wrong.
Back when animations were done by hand, Disney's animators handled this by rotoscoping (drawing over...) live action, which created very convincing results even when the movements were exaggerated for effect.
Doing the same with CGI without a template is an unbelievably difficult challenge. It's much easier to create cartoonish exaggerations than to get spot-on perfect realism.