This is all good, but how does one store, index and efficiently search for near matches?
I have a side project with terabytes of photo URLs, and hashing&indexing them has been an itch I could not scratch.
I stumbled upon metric trees for nearest neighbor queries in metric spaces some months ago and wrote down notes here:
http://daniel-j-h.github.io/post/nearest-neighbors-in-metric...
Hope that helps.