Humans work by abstracting concepts in what they see, even when looking at the work of others. Even individuals with photographic memories mentally abstract things like lighting, body kinetics, musculature, color theory, etc and produce new work based on those abstractions rather than directly copying original work (unless the artist is intentionally plagiarizing). As a result, all new works produced by humans will have a certain degree of originality to them, regardless of influences due to differences in perception, mental abstraction processes, and life experiences among other factors. Furthermore, humans can produce art without any external instruction or input… give a 5 year old that’s never been exposed to art and hasn’t been shown how to make art a box of crayons and it’s a matter of time before they start drawing.
ML models are closer to highly advanced collage makers that take known images and blend them together in a way that’s convincing at first glance, which is why it’s not uncommon to see elements lifted directly from training data in the images they produce. They do not abstract the same way and by definition cannot produce anything that’s not a blend of training data. Give them no data and they cannot produce anything.
It’s absolutely erroneous to compare them to humans, and I believe it will continue to be so until ML models evolve into something closer to AGI which can e.g. produce stylized work with nothing but photographic input that it’s gathered in a robot body and artistic experimentation.