My hypothesis is that we will first use visual recognition for a few highly relevant fields. But after that we'll expand to almost all objects.
The reason I ask is because if this is not a pressing need for you, it might not be a pressing need for other people. And even if it is a pressing need for some other people, it is only for a few not for most people. Good luck.