Meet the algorithm that can learn “everything about anything”
gigaom.com
gigaom.com
http://levan.cs.washington.edu/ngrams/objectNgrams_cvpr14.pd...
So basically there are two categories of "learning" involved in this sort of research, supervised and unsupervised. In supervised learning, someone gives the computer a long list of concepts and their attributes ("frog", "green frog", "jumping frog") and a set of pictures to go with each item, and feeds them into a visual-recognition algorithm. In unsupervised learning, the computer is given a concept like "frog" but then has to discover all the variations itself and get its own visual data to match.
The claim in this paper is that they have made the unsupervised learning as strong as the supervised learning. That is, they give the computer a concept ("frog"), it goes and searches through Google Books for common variations ("green frog", "jumping frog") and then uses Google image search to fetch images for each of those queries. They can then remove the obvious false positives (they test to see which images seem to screw up their learning algorithm and leave those out), and the result they get is on par with the supervised learning methods.
----------------------
In my opinion, this is only mildly interesting because Google Image Search functions based on human input anyway -- Google knows the difference between a "frog" and a "jumping frog" or even a "camel" simply because people on the internet caption such images and Google can make associations between images and their captions. Essentially, what the researchers have managed to do is outsource the work of some grad student to millions of people around the world through Google.
Of course, it could be argued that there is some sort of parallel with what humans actually do (we know what things are called because we hear other people call them that), but even if I didn't know the name of an animal I could still tell you when the same animal is in different pictures, and I can also tell you when it's jumping and what colour it is. I don't need to have someone caption the image for me to understand the broad range of situations to which the caption "jump" applies.
I wonder if this has anything to do with the fact that we can jump too. That we can translate the frog's position into something we do as well.
of course one can argue that we can do the same for non-anthro-moprhic things as well. What i think is that we dont directly relate pictures, as the software is taught. What we do is translate that 2D picture into something we'd see in the 3d world. And that 3d "vision" isn't just another image. It represents an object in our world. something that has shape, existence etc. something which we can observe from other senses as well. For us a picture doesn't always represent an abstract thing, an arbitary pattern of colours. It usually represents something concrete. Something about which we have tons of other pieces of knowledge as well.
So we relate pictures by checking if they map to the same real-world object. And here that "object" is a sort of nexus of many pieces of information we have on it which is a product of many direct and indirect human experiences.
So i don't really think that we are in a position to teach a computer to do anything like that.
I'm more interested in metaphor and analogy.
My 3.5 year old son said "look at the rain! It is bouncing like hopping frogs!"
I don't know if he created that. It's not in any of his books. I guess he jumps like a hopping frog at nursery and transferred that to rain.
I'm not so interested in a computer that is trained on frogs, and which sees a hopping frog and describes it as such. If it saw a hopping cat and said this thing is hopping but I don't know what it is, then I'd be interested.
Am I being too harsh on the robots?
Another way is to use some kind of optimization to find an input which produces that pattern (e.g. backprop to the pixels themselves.) This will give you the image that most strongly triggers that output. Not necessarily a typical example.
None of them seem to work towards a fundamental understanding of what the concepts mean. I'm not sure how that could be accomplished though.
Is there even a meaningful distinction to be made between beeing able to identify a concept and understanding what a concept actually means?
IMO "Imagination" itself isn't intelligence but an input to the last "intelligent" activity.
Computers can be taught to exhibit different "intelligent" activities and that is AI & Machine learning is all about.
Human imagination works without any knowledge skill or learning (it is just my understanding, and i might be wrong).
Producing rough and wild-west imagination isn't problem of AI.
Analogous to the philosophical zombie thought experiment, I think that "real intelligence/understanding" is indistinguishable from simply being able to perform actions to accomplish the same tasks that humans traditionally consider to require intelligence.
Of course, that's one of the PR problems that AI has always had: once a computer can outperform a human at some task, that task is no longer considered to be something that requires "true intelligence." Most people would consider someone who can multiply numbers together to be intelligent, but when computers do that (incomprehensibly faster and more reliably), few people consider even for a moment that it's AI. Same with more advanced mathematics, like computer algebra and automated theorem proving. Same with facial and voice recognition. And I'm sure it will be the same with self-driving cars.
Until I can just talk to a computer and it always understands what I'm asking, there's no AI is there? It's a fuzzy concept to exactly define, but it's pretty obvious what we all mean when we talk about it.
Brute forcing billions of moves of chess isn't going to achieve that, so it's nothing more than a cute trick that lay people don't understand is a trick so we have to tell them it is.
David Copperfield can "fly" but he can't just fly on demand, he can only fly in very special circumstances in a tightly controlled environment.
He's no more a flying man than deep blue is a thinking computer.
Building new concepts from raw inputs is the difficult part at which humans are better, but now computers are showing that they may be able to do that too.
There are people who are looking at broader ideas of "what is intelligence", sometimes even the same people doing the more statistical research, but it's an open ended problem.
The main reason science/math/statistics doesn't spend time worrying about philosophy is because you can't really specify the problem cleanly - and having a clear idea of the problem is 90% of the way to finding the solution.
[edit]: If you want brain-inspired ML stuff, it's worth looking at what neuroscientists and cognitive scientists are doing. From what I've seen of some of their work they're developing ML algos as minimal tests for understanding how parts of the brain or consciousness works.
It would be impossible completely describe a concept such as "horse" to a person with no knowledge of animal physiology, without recourse to pointing at a a horse or impersonating a horse noise. Resorting to analogies with other concepts - "a horse is like a cow" - doesn't work because the person has no knowledge of the other concepts. This person could be a young child or an untrained computer, and they can be taught fundamental meanings of concepts by feeding them information in the form of sights, sounds, smells etc.
I haven't thought about this very deeply but these are the thoughts that always come into my mind when this topic comes up. Concepts have no meaning without relating them to sensory input. Humans learn by connecting the two. Why can the same not be applied to computers?
That said, it doesn't at all mean you can't learn anything meaningful from it. For example, you build a machine learning system to predict the missing words in a sentence. "A horse is similar to other farm animals like ____." Machines are getting better at this kind of thing, though still far from human level. Google's word2vec for example can take a word like "horse" and list the words that it is most similar to. You can subtract the representation for "man" from "king", add "woman" and it outputs "queen".
How would you define “understanding” and “actually means” in your question?
[1] http://gregegan.customer.netspace.net.au/DIASPORA/01/Orphano...
fwiw, that "Webly-Supervised Visual Concept Learning" reminds me of the stuff that Hinton et al. do re: unsupervised (concept, etc.) learning (using restricted Boltzmann machines, and so on.) Good talk on the subject (of deep learning, etc.): https://www.youtube.com/watch?v=AyzOUbkUf3M
[1]: read online here: http://bookre.org/reader?file=222997
Also, about the LEVAN thing... given the amount of data available online, both in various structured formats and unstructured formats, don't be surprised if deep learning will yield better and better results moving forward. To me though, they mostly seem evolutionary rather than revolutionary. I mean if you look back at the AI field, during the days before the "AI winter" came, huge amounts of data is one thing researchers back then didn't have available. This is not to say that there haven't been advances in learning algorithms at all recently. ..
And there is what one group chooses to do with a very special neural mod...
In defense of the LEVAN thing, though, I didn't see any claims that this is science at all, more like an exploratory application illustrating an algorithm.
Deep Learning is actually worth the hype though. It has 2 main merits that are interesting.
1. Auto Trend Discovery
2. Plays very well with parallelism
The main problem, which I'm hoping to fix, is feasibility and ease of use. Neural nets to the untrained eye can be a black box that takes a really long time to train with little to no reward.
The hype isn't all for naught either. I'll elaborate if asked, but won't bore you guys otherwise.
The results coming out from different tasks are currently blowing away many of the old school algorithms in tasks like sentiment analysis, speech to text, object recognition, among others.
The overall idea is a support, services, training model.
Support: Install/Infrastructure
Services: Help with data etl/onboarding, tuning
Training: How to use deep learning, how to think about it, and how to run/setup clusters. This will be run through the machine learning academy I teach at[1]
My point is that mixing business goals and scientific truth is dangerous if not handled carefully. That said all the best with both your goals.
Marketing hype and machine learning really does make things convoluted. That being said, what DOESN'T the press exaggerate?
The track record for AI has definitely been over promise and under deliver. I think the hardware is getting there where we can start making real progress though.
In this case, google is putting its money where its mouth is.
Thanks for the wishes, it's been a great ride so far. My overall goal with this is to bring deep learning to everyone else. I see real merits in auto discovery of trends to automate some of the worst parts of machine learning and would like others to see these benefits as well. Seeing cool apps built with it is another side goal of this as well.
I thought that this was an early, encouraging sign that the models being used were fundamentally valid. It's entirely consistent with the track record for natural intelligence.
Even the emergence of neural nets in the past few years has been due to hardware increases.
Still needs a bit of work in terms of linearity, but running it on a 20 node cluster was easy. Not going to benchmark it more than that right now, but it's really starting to come to life.
In particular -
1) what does 'automatic trend discovery' mean? We've been able to do change-point detection, linear regression, etc. for hundreds of years. If you're talking about automatically learning a feature representation, then there are other algorithms that can do this, in a much simpler way. If you're arguing that it produces better representations, then make that argument.
2) This is almost _completely_ false, and indicates a substantial lack of experience of ML beyond deep learning. Other machine learning algorithms (SVMs, LR, even some decision tree algorithms) are _much_ easier to train in parallel - this is (partially) because your objective function has certain nice properties that allow you to combine partial solutions together that are produced in parallel (convexity, separability). When you're using gradient-based methods on an incredibly ugly non-convex function from a multi-layer neural network, you're in a completely different world.
Granted, there have been techniques coming out for training in multiple address spaces, but these are _hacks_ to get around the ugly structure of the problem, not the principled approaches that exist for other algorithms.
I don't have any perspective on your deeplearning4j library, but I'm skeptical of its utility given the existence of existing well-tested deep learning libraries written/contributed to by renowned experts in the field (e.g. cuda-convnet, caffe, torch). This stuff is a) very tricky to get right, b) very tricky to debug, and c) very performance sensitive. Just a quick pass through shows zero references to CUDA/GPGPU programming, so I'd suspect performance is going to be significantly worse than the aforementioned libraries.
Yes, I am talking about learning better representations. See hinton's deep autoencoder work as a prime example of this comparing PCA to RBM based methods for topic detection[2].
2. Google and people way smarter than I am seem to be doing just fine with this[3]. That being said, I didn't say that random forest (with whole companies built on this parallelism[3]) or any of the algorithms WEREN'T friendly. I would say one of the main appeals for deep learning is the scale of data with which it can benefit from.
Feel free to be skeptical all you want, if the researchers want to take the time to write a full stack distributed framework, I welcome others in to the game. The problem with the packages out there right now, (being matlab, python) are training times, and integrating in to an actual ecosystem. I'm addressing this this year at 2 different talks[5][6].
Replying to your last point, I use blas underneath for all of the matrix calculations, I will be adding GPUs later this year, and yes you're right,this stuff is hard to make. I also wouldn't be publicizing it if I wasn't already using it in production applications. Frankly right now though, I use cpu matrices right now, because I can fire this up on AWS (without the limit of GPU RAM), and it's practical for hadoop deployments. Honestly whether we like it or not, GPUs take a lot to get right. NVIDIA[7] and AMD[8] are going to make my job pretty easy though.
To end, if I was afraid of every little obstacle, why do anything in the first place? While you're hiding behind a throw away account, I'm actually trying to put this in the hands of people who don't have the time to learn every little thing about neural networks. At the end of the day, I follow the papers very closely and enjoy what I do. I also work on all sorts of different techniques for different problems combining different machine learning algorithms for different tasks (just like anyone else would). This framework is my way of getting this out to everyone else. If you have a deep learning framework, I'd love to see it, maybe I could learn a thing or 2.
[1]: http://zipfianacademy.com/
[2]: http://www.cs.toronto.edu/~fritz/absps/esann-deep-final.pdf
[3]: https://bigml.com/
[4]: http://static.googleusercontent.com/media/research.google.co...
[5]: http://hadoopsummit.org/san-jose/schedule/
[6]: http://www.oscon.com/oscon2014/public/schedule/detail/33709
[7]: http://www.jcuda.org/jcuda/jcublas/JCublas.html [8]: http://developer.amd.com/tools-and-sdks/heterogeneous-comput...
Edit: In case that seems foreboding at all, what I mean is: this looks really cool and will probably make an awesome Show HN, whenever you think it's ready for a post of its own.
[patent pending]
See http://levan.cs.washington.edu/?state=show_aboutI am not saying that i agree with him, i am just trying to clarify his point
* Grant applications and future work sections excluded :)
Usually it's just weak journalism rather than researchers trying to overinflate.
That said, given that Google et al. are scraping every corner of the bottom of the barrel in search of advances in search, I think the good times for AI peeps will continue for some time...
Okay, let the thing 'learn' about the Kuhn-Tucker conditions by searching on Google and reading, say, Wikipedia or some books at Google or Amazon. Then have the thing show that for problems in functional form the Zangwill and Kuhn-Tucker constraint qualifications are independent. Do that and I will start to believe that the terminology 'deep learning' is appropriate. I'm not holding my breath.
Yes, it may be that in some rough sense the kind of 'learning' it is doing is roughly like some of the learning of a child of, say, 2 as it is starting to learn about language and things. Yes, it may be that such 'learning' is a significant part of the intelligence of, say, a child of 3-5. Maybe. Big, huge maybe.
When I was working in AI, I noticed the terminology had been cooked up to imply much more than was being accomplished. Now, as I understand it, there is a specific definition for the current AI term 'deep learning' and has to do with the 'depth' of where adjust parameters in a neural network, not how 'deep' the 'learning' is about the subject in question. Cute terminology.
I quote the readme:
> This is an implementation of the "Learning Everything about Anything" system. The system is implemented in MATLAB, with various helper functions written in Shell, Python, MEX C++ for efficiency reasons. For details about the method, please see [1].
This readme contains instructions on using the code, as well as accessing/using already trained models for various concepts.
For questions concerning the code please contact Santosh Divvala (http://homes.cs.washington.edu/~santosh) at santosh@cs.washington.edu.
The software has been tested on Linux using MATLAB versions R2011a. There may be compatibility issues with older versions of MATLAB. At least 4GB of memory (plus an additional 0.75GB for each parallel matlab worker) is assumed.