If you reverse the output of a CNN "hand" classification, it'll give you images that resemble the geometry and shading of fingers, palms, nails, knuckles, etc. -- these, I submit, are the distillation pattern matching for the actuality of "hands". Under no circumstances will it give you the five widely-separated fingers which a child draws. That's because the child-drawn hand is based on literal visual stimuli, but rather on an abstract logical model of a hand. That logical model is fully integrated with a similarly abstract model of the world, and includes functional relationships between abstractions, like the knowledge that "hands" can open "jars". The value of these being logical models rather than matched patterns is that they can then be extended to include never-before-seen objects. Confronted with a strange but roughly jar-sized object, a child can surmise that maybe it, too, can be opened with hands. That isn't pattern-matching: it's algebra.