Reverse image search engine
tineye.com
tineye.com
I think we're going to see some very interesting developments along these lines very soon. Scary stuff too. Imagine submitting a picture of yourself and finding out what the internet knows about you based on your physical appearance. Better keep those Facebook profiles private, folks! More than that, you'll have to convince your friends to keep their profiles private if they have pics of you as well!
<tangent> This is what is rather frightening about the next web; even if you want to remain anonymous, you're going to have to do battle with all the other folks who are more than happy to post and tag pictures of you for the world to see (with good -natured intentions, I might add). Remember that embarrassing moment at that party where you had a little too much to drink? Oh, you were too drunk to recall? Well, it's on somebody's public Facebook profile now. With your name on it. And if I am your employer, what's to stop me from taking your badge photo and plugging it into a service to pull down other pictures of you from the cloud? :O </tangent>
Anyway, back to the matter at hand! I do see your service as being particularly valuable to IP holders who want to know who is displaying their copyrighted images or logos without authorization. If your site were comprehensive enough, you could probably go freemium and become a paid tattle-tale. Take that a step further and "For a nominal fee, you can click here to have our partners at LegalZoom.com send a takedown notice."
:)
Flickr is more public, but they don't do any user-graph linking at all, the photo annotations are plain text. I don't think the annotations are even exposed to google.
Example:
http://imageheader.com/alpha/index.php?url=http://farm4.stat...
Actually I think this sort of technology would be just as useful for the dating sites to help weed out potential fake profiles.
Your TinEye approach, later that day, via email: "Hi. How's it going? You don't know me, but I surreptitiously took your photo in the park earlier today, uploaded it to a reverse image search engine, and eventually uncovered your personal data after an exhaustive web search. Care to get some coffee sometime?"
Why don't you A/B test these two approaches to find out which works best?
Your Mechanical Turk mission, should you choose to accept it, is to identify this porn actress...
The algorithm, however, is beyond awesome. They found half a dozen instances of my screenshot on the Internet, including my site, some download sites, and a Chinese pirate or two who had gone to the trouble of watermarking my image.
I think Getty just had kittens.
For some cool examples, check out http://tineye.com/cool_searches
Also, I recommend that you click the "compare images link" under each result image after you perform a search, to see which part it matched.
I've used tineye several times, and it's found the sources of heavily photoshopped images before. They've done a great job.
TinEye was created by Idée Inc. Idée develops advanced
image identification and visual search software for photo
wire agencies, stock photography firms, entertainment media
companies and some of the world's leading imaging firms
including Adobe Systems Inc.
In other words, yes they intend for it to be used by content owners to find unauthorized use of their IP on the web. On the other hand, they claim to respect robots.txt and give their crawler name on the same page.if you browse the results, it even finds the image's use in formatted/manipulated graphics
Search is a service built on matching and indexing.
For each image in the index, break it up into 4x4 tiles, then store a hash code for each tile. Then repeat the process, but offset the boundary of each tile by 1px along the X axis. Repeat 2 more times. Then, for each offset along the X axis, offset down along the Y axis. So you store 16 hashes per 4x4 pixel area.
Now, when someone searches for an image, repeat that hashing algorithm for the source image. The results page then returns any image that contains a 4x4 tile that is also contained in the source image, ranked by the number of tiles within the image that is common between the source and result image.
The end result is that you can see how the features within an image are used in other images -- so if someone takes the red stapler from Office Space ( http://www.yunasville.com/img/102005/milton.jpg ) and puts it into a different image, and you search for that red stapler, the results page will still return the photoshopped image, because it'll match the 4x4 tiles on the stapler in both images.
I've explained this in a convoluted way, but hopefully I've communicated the essence of the idea.
On one hand, there will be more results to filter through, and it's more computationally expensive. But that's fine, the image results are still ranked effectively. On the other hand, it's more computationally expensive.
First, looking at the examples, it doesn't seem to be restricted to exact images, not by a long shot.
Second, your approach is hopelessly naive. ;) Consider: rescaling, re-encoding using a lossy image format, color adjustments, and so on.
Good news though, I bet your approach is way, way, _less_ computationally expensive than whatever it is they are doing ;)
(My guess would be something like SIFT or SURF, probably minus (some) rotation invariance to speed things up, combined with a whole bunch of hackery to make the feature database search suitably fast while still acceptably accurate. Worth your time to google + read up as those algorithms can achieve positive red stapler identification fairly robustly, but you'll be in for rather a lot of math.)
So it's not what some of us might have been afraid of.
Very impressive work. I can't think of any real need for it myself, but it is cool.
But will people pay for such a service?
There are just so many companies starting with seemingly no way to make money.
It's not a stretch to say they will pay to find stolen images. I'd venture they'd pay pretty well too.