Show HN: Visual Search using features extracted from Tensor Flow inception model
github.com
github.com
In future I plan to add more images ~2 Million in the same domain. Test various combination of nearest neighbor indexes and multiple vectors per images using some form of a multibox style detector. I will also add a script to launch spot GPU instances via Cloudformation to economically index images using S3 and SQS. I am building a companion iOS swift app however, since Tensor Flow hasn't been ported to iOS yet, its still in development.
https://engineering.pinterest.com/blog/building-scalable-mac...
I beleive that most of the strength of VGG (and Inception) vs Alexnet is that it was able to learn the feature relationships better, not because it learnt better features.
VGG is pretty computational intensive, which is why Google concentrated so much on the computational complexity for GoogLeNet/Inception/ReCeption
So if you are just using the features directly then it would make sense to use whatever was the quickest to compute.
I am interested in assessing if there are any tricks that could be used when querying from a mobile device. In such cases feature extraction can be performed on the device itself, with only feature vectors sent over the network. In case of pinterest, another special case is that a lot queries are performed on images already present in the system. The user simply readjusts the bounding box to highlight the object of interest. In this case they can simply pre-compute 4~20 crops per image. Online feature computation is much more expensive / complicated than offline.
But it would be interesting to know if that is better. I'd imagine most phones have some kind of hardware support for resizing images, so it might be better to take advantage of that and then do feature extraction on a server?
AlexNet will run at ~10 FPS, I guess.