Very cool. (1) Have you thought about using HOG templates for the matching step? Seems like you might get some interesting structural encoding similarities there. (2) When generating image patches, could they be different sizes? Or is the idea to make a simple grid?