"one is painted by a human and another one is generated by artificial intelligence based on a photo and a style of a painter."
I find this to be easily "beatable" by simply judging which one is most likely to have its origin in a photo (vanity, and such) with a fallback on the one with a lot of repeating patterns.
I don't think programs written to imitate a craft, or even to learn to imitate a craft from examples, count as AI, no matter how impressive.
I disagree - nobody would say drawings/paintings (by humans) based on photos or observations by eye are less legitimate than scenes completely imagined.
Interestingly, this is a Turing test that, when applied purely to humans, doesn't require said humans to share a common language.
I have to quibble though that the Turing test is meant to be a test of intelligence, and this sort of task seems pretty different, although it may actually be visual-AI-hard (to invent a new term). So may deserve a different name.
Those photo modifications kind of fit into that theme: if you can't distinguish the filters from real art, they passed an indistinguishability test.