Given the amount of training image data with hands (and text in the image), I don’t understand the lack of specific detail. For hands it’s even stranger - the models add fingers, for example, which seems like something that the training data never sees.