TinEye: image recognition search
tineye.com
tineye.com
http://news.ycombinator.com/item?id=535818
http://news.ycombinator.com/item?id=423679
http://news.ycombinator.com/item?id=356971
http://news.ycombinator.com/item?id=302020
(and several more articles discussing the company)
Even with a small index, though, this is pretty damn cool.
It is pretty impressive.
Unfortunately, they might have to add a Twitter suffix to their name to attract funding.
edit: ok I'm watching the video where you, or someone, is explaining the service. I was wondering how limiting is the access to data? I'm under the impression that you crawl the web no matter what and clients piggyback on the crawl stream and do analytics on it. So, who makes calls on web crawling method, I guess you? What if a client wants to crawl and analyze only a specific domain, country specific, for example? And how often and what exactly gets crawled? Lets say a client wants to implement a news.google.com or google alerts (or even tineye) - just as an example - it would analyze a web crawl stream from you and get data out of your system to their servers for utilization? How would such a, presumably large, data get transferred over to the client? What is provided as crawled data? Only a html stream, or the whole page that includes js, css and pictures? Or would a client need to get image links from your crawl and get images themselves? Sorry for lots of questions and incoherency :) but it does look interesting.
1. Who makes the crawling method? We give you the ability to write your own crawling logic. You don't have to though.. we have a default crawler that runs as well.
2. Crawling specific pages? You can specify pages to crawl using regular expressions or your own custom code.
3. How often and what gets crawled? Up to you :)
4. How does the data get transferred over? It's better if you don't transfer all the data over. You can push in your own compute functions to process the data you crawl, and just return much smaller result data sets.
5. What's provided as crawl data? You can specify the results you return in your code. It's up to you.
If crawling logic can be at clients control, but there is a provided crawler that runs well too, and those prices, it sounds like an excellent product you've got there :) Let's just go over this tineye example we have here - they need to crawl web and retrieve images - so we have "Only pay $2 per million pages crawled" where they would pay $2 per million pages crawled, but what would they pay for retrieving images then?
I see so much potential in this service, now that is one hell of a product there at one hell of a price - congratulations on it, I hope you do well with it!
Any computation done on the images would be priced at $0.03 per CPU-hour.
http://labs.ideeinc.com/visual/
On this page, you have to select the query image from a set they provide.
On the page linked below, you can upload your own.
I tested both my company logos in it, and nothing came up. This means they don't infringe. :-)
Countdown until they're bought out by Google or similar? ...