Andrew Ng predicts half of web searches will soon be speech and images
venturebeat.com
venturebeat.com
How about framing the problem as proactive vs reactive searches? Piece together enough fragmented data about me to know what song I'll have stuck in my head and don't know the name of, recognize my email contains a course syllabus and auto-populate my calendar with assignments and study times... a million other proactive tasks all done with me in mind.
Geez these guys are talking about changing the interface... try and get rid of the interface all together!
Google/Bing have done well with keyword/phrase searching with results sorted by popularity and date.
Ng seems to be thinking that the big change will be having speech and images coming from the users as their input to the search process. My guess is that this will not be very important. I also guess that search for Internet content based on speech and images will become more important.
Yes, likely Ng can get a lot of pictures of what he knows to be, say, Ferraris and use some of them as the training set with a neural network to identify Ferraris and test the training with the rest of the pictures. Okay. Maybe his neural network will be able to identify Ferraris. So, he could repeat this training for, say, 100,000 objects -- Fords, bread, airplanes, jewelery, Victorian houses, .... Maybe there will be some value there.
My view is that the future of Internet search is quite different.
Ng wants neural networks to identify things, maybe Ferraris. Then he wants a search user to send a picture as their input for the search they want to do. So, then the user might be able to find more pictures of Ferraris. Maybe.
For what Ng is doing, I doubt that flowers would work because there are far too many too different cases of flowers.
Search by keywords/phrases by Google/Bing has worked very well for a huge collection of Internet content.
But I am guessing that in a sense Ng is correct about images and sounds -- there stands to be a lot more such content on the Internet in the future.
Generally I'm guessing that there is also a huge collection of Internet content, searches people want to do, and results they want to find where search via keywords/phrases such as via Google/Bing is from poor down to useless. Thus, my guess is that a new means of search is needed.
What do you think?
If you have enough domain experience to search for "continuous rotation rotary intermittent rotary" then you'll find it, but if you are mechanically illiterate you may not know what intermittent means or rotary... maybe.
It would be a truly amazing display of AI to be given a really poor sketch of a Geneva mechanism, it'll find a really nice blueprint. I'd be impressed if this found the hypoid gear in a differential. I'd be more impressed if someone who doesn't understand the concept or reason for a torsen differential was able to none the less search it.
As a concrete example theres a pretty impressive lego torsen(-ish) diff out there. Its easy to find if you google for the terms. I'd be impressed if you could give a sketch to a search engine and find this lego diff.
For the searches you mention, I believe that maybe for one of them it could be possible to improve on Google/Bing, but I don't believe that real AI would be needed. For describing how to build such a search engine, that might take more than the 10,000 character limit on HN posts!
Or, there's a lot of content on the Internet; a lot of it, a lot of people want; to get it they want to do searches that promise to be able to find that content. Nothing more obscure than that.
Only if you see it as a supervised learning problem. An alternative is to find the nearest matches, after which you can let human intelligence make the final visual match. Often, the webpage containing the matched image will have enough context to identify the object being searched for.
Or is a 'search' something that a person dictates to a computer to answer a question?
But if they could do accurate speech recognition with a very quiet whisper (maybe using lip reading technology too) then I could see it completely dominating text searches in usage.
Text tags on images have a limit to the accuracy of the results.
My wife and I usually use speech input for web search on our phones. Also, the current Google image search is very nice. Have you tried using it? Go to Google image search and drag a photo to the text input field. I have used this to identify pictures of small mechanical parts and also to identify plants.
A year ago I showed my kid (then 5) Google voice search on iPad. Her eyes lit up and off she went on Youtube and later Google search itself... First obscure animals - obscure to me, but apparently mentioned on 'Wild Kratz' and other animal shows - Then stuff she heard about on'Magic Schoolbus' and Brainpop.
Judging from what she's been showing me, it works for song titles and lyrics as well, and she loves the voice search interface on our Amazon FireTV as well.
As a result, I find myself using it more often. I love it on the FireTV. If I had a single button for picture search on my Sammy, like the Amazon Fire Phone has, I'd use that too.
I will show her picture search tomorrow. Can't believe I forgot that.
Edit: after reading the article, one more important group comes to mind: those for whom English is not their first language. Calling my parents tomorrow.
A friend once quipped, "the best porn site in the world is Google." As "reverse image searches" (where you provide an image for the service to find image matches, similar photos, associated information on the photo to figure out the identity of the person in the photo) improve in quality and data mining from the pages they find the related images from, the more people will shift their behavior to this from the current text based searching.
There will of course be searches for similar looking wardrobe or other commerce or travel reasons. But online porn is a behemoth that will overshadow any other image based search motive. It is driven by one of man's three basic desires after all (nourishment, sleep, and sex).
Especially when we become so "plugged in" that I can search for what I see. Right now, if I want to see "tomato recipes" I'm probably not inclined to pull my phone out, take a picture (and who wants that on their stream anyway?) and then paste picture + recipes? I'd rather just type "tomato recipes". But if I can search what I see and speak, e.g. "oh look, tomatoes? Find recipes!" that's much easier.
For produce there are even more things I might be interested in. Is that a good tomato or what is that deformed looking red plant?
Essentially I'm saying that image searches have just never been useful for me without supplying textual context. Speech has the potential to help with that but just doesn't seem to work well in consumer devices yet. Speech will get there but image searches just aren't useful without speech or text.
This way you wouldnt have to do that.
I bet you've already done something like this with QR Code enabled posters.
And then at some point, it'll become a live-streaming thing. i.e. instead of actually flipping out your phone and taking a picture, it'll be some google glass type thing. You walk past a store and you look at it and it's recognized instantly. You can then get the wiki page, the floor plan from gov records, its online store and opening times from various online repositories of such information. Of course this is already possible using GPS, just an example.
I think mobile image search will see a big uptick if there's a good interface for it.
- take a photo of a street scene and ask, "how do get home?"
- or more generally, "where am I?"
- If I hate reading every ingredient on a menu to check for allergies, why don't I point my phone at the menu and have it tell me what I would like best?
- "what's in this sandwich?"
- take a photo of water, "is this safe to drink?" (maybe realistic with near-IR?)
- "what plant is this?"
- record birdsong, "what species is that?"
- photo of text fragment, "who wrote this and which book is it from?"