These models don't understand relationships between objects in a scene, especially between distant objects. So they can't do hands for the same reason they can't get legs on a table right. They know roughly what a table and a table leg look like, but they don't understand that there needs to be 3-4 of them at least, and they need to be spaced so that the table sits level, and the perspective they should have as a result. So, I've seen tables where it kind of gets it right that the legs are in the corners but then as the table legs go down, the front ones are mysteriously behind something that ought to be under the table. And sometimes it kind of loses track of a table leg or two - they melt into the background.
Very similar problem with hands. They need a very specific orientation and shape and the fingers all need to consistently point in the right direction, and typically the same direction (except for when they don't like with a pointed finger, etc).
Curious as to how these models handle it so much better than prior generations. Is it something novel, or a specific hand-based fix they put it, or is it just "we made the model bigger"?
The fingers themselves are also almost identical, but not really. If you learn a "platonic finger" it's not good enough, you should learn each finger individually. There is only so much you can spend on them, you got a million other things to learn. And the raters of the model are much more likely to penalize a bad face than some off details in a hand.