Another reason I saw was that models were trained on 512x512 "portrait" images including very few hands. Added to the inherent complexity of hands, this throw off their generation.
Humans seem terrible at it in very different ways, and definitely don't get as good at other parts before getting good at hands.