I was thinking along those lines, but don't you still make it easier by giving away some details about Waldo's facial expression, the relative size of his head, etc.? Still a pretty illustrative analogy though.
How about get the coordinates of Waldo's position in the image and hash them?