But you are not capable of reproducing from your mind something that is in that artists' style, after that viewing. If you show someone a painting and ask them to recreate it, even setting aside the skill gap, you do not get the same painting back. You get a different painting, with some overlap, with the "focus" of it being what the person paid the most attention to during their viewing.
In this way, ML training is just not the same as a person being inspired by or even being asked to recreate another person's creative work. Over the course of making something, an artists' "voice" would be best characterized I feel as the tiny choices they make along the way that all point to and reinforce a larger point or purpose to the piece. This "voice" shows up in all creative output, not just spoken word. That is what I feel people are feeling is lacking in generated art: because a machine-learning model does not have a voice, it does not have intent, it has a mandate from a third party from which it tries to draw from, and instead of making numerous, tiny but contributory choices, it instead decides on a weighted average of all the choices made in the art that the model was trained upon, which is simply not the same thing.
It makes all generated art have this very sterile, soulless feeling to it because these tiny choices that would otherwise be made by a person trying to illicit an effect are instead just the machine sort of shrugging and being like "well in most things I've seen where a woman is sitting this way, her hand is tilted this way" but it doesn't know why the hand is tilted or what that means for the subject, which means the hand-tilt might be applied to subjects for whom it makes absolutely no sense at all to tilt the hand.