Create an algorithm to distinguish dogs from cats
kaggle.com
kaggle.com
If the animal comes, it's a dog. If it continues without looking at you, it's a cat.
/joke session
And then read the comments and detect indicative words or expressions connect with each animal
if (human.feedanimal() == true) { animal.type = "Dog"; human.name = "GOD"; }else{ animal.type = "Cat"; animal.name = "GOD"; }
Incorrect. The cat doesn't depend on your validation of its status; if you feed it, you just increase the chance that it thinks you are a subject worthy of its time and attention.
EDIT: If anyone one thinks we can start working on this now, I'm game.
EDIT: I would definitely be interested in building something like this. iOS/mobile app? I have basic experience in ML and have written an ANN in C++ to classify letters (they were 'pixelated' images, 1 and 0's).
Bloody ML Summer...
You can take photos of leaves and it'll identify them for you :)
It's not immediately obvious to me how useful such an app would be btw. Unless I of course misunderstood what a "real life pokedex app" is :).
Is state-of-the art for that kind of recognition deep learning?
The kaggle competition is for determining if it is a dog or cat, so it's a bit unlikely that one of these approaches would directly work (although they might be adaptable to the task). See my other comment [2] for a lighter-weight approach that is likely to do just as well, if not better.
In computer vision, these two types of images are traditionally handled separately. First, a detector for a class (like "dog" or "cat") is run across the image at all locations and multiple scales to find where the things are. Once you have the locations, then an image classification algorithm is run for each detection window to either confirm it, or to give you more information about the object.
The latter often takes the form of giving more fine-grained category information, such as what species of dog/cat it is. Both leafsnap [1] and dogsnap [2] take the form of this type of program; i.e., they both assume that you've captured a single subject, roughly centered in the photo window, and that you already know that it's a plant/dog.
Sometimes you don't have to run a detector even if the object is not the focus of the image, if the context/setting can narrow down the answer for you. For example, if you were deciding between dogs and airplanes, it would be pretty unlikely to see a dog on a runway or a plane in a living room, so just by classifying the entire image, you can do reasonably well. That's not the case here, as dogs and cats will, for the most part, appear in pretty similar environments.
So if I were attacking this problem, I'd first see how many images were of the non-focused type. If not many, I'd basically ignore them and focus on building a classification system. Note also that if you're constrained to make a hard choice between only two classes, that's a much easier problem than a more open-ended "what is this?"
As many have pointed out, deep learning approaches seem to be the current state of the art on classification tasks such as these. But deep learning requires a lot of training data to be effective. A procedure I've been hearing many people use to great success is to use the Imagenet [3] hierarchy and images to train a deep learning classifier (i.e., as if you were going to compete in the Imagenet Large Scale Visual Recognition Challenge [4]). Then use the trained network, chop off the last stage (which makes the final prediction), and replace it with an SVM trained on your specific training data. In this way, you'd be using the network only as a feature extractor.
I'm happy to try and answer other questions.
[1] http://leafsnap.com or see my project page for more details on how it works: http://homes.cs.washington.edu/~neeraj/projects/leafsnap/
If one could locate the face in the test set, she could also presumably find some landmarks of interest: eyes, nose, mouth, etc. Considering that dogs typically have longer snouts, cats have pointier ears, etc, this data could be used to differentiate between a dog and a cat. There would be difficulty dealing with awkward angles and bad lighting though.
[0] http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.6.35...
[1] http://www.eng.auburn.edu/~troppel/internal/sparc/TourBot/To...
Standard computer vision features like HOG (histograms of oriented gradients) or SIFT will probably do much better, or the deep learning features others have mentioned.
Your larger point of adapting a face detectors for animal use is well taken, though probably overkill for simply saying "dog" or "cat". You need that level of detail to identify which breed (e.g., this is the approach that dogsnap takes), but not for the base distinction.
The other way to go would be to train a deformable parts model (DPM) detector [1] for dogs and cats. DPMs are the current state of the art in detecting objects, e.g. as measured on the pascal VOC benchmark [2].
or
"You too could solve this problem, a get a Phd and joined that overcrowded labor market"
Just consider that if you have M categories and you have N Phd students who can each four years to create one clever algorithms to distinguish category i from category j, then you need M(M-1) Phd students for a complete classification system - which when you consider many, many categories there are in human knowledge, works out to being more than can even be pumped out by excess student loans today and exponentially more than can find tenured positions.
IE, once you'd add to the "deep but not wide" algorithms of computer vision, And twenty years ago, we might have believed this adding-to would lead to something broad and general but it's been twenty years and the trend is becoming clear.
See:
To quote @BigDataBorat (Twitter): 90% of data is unstructure. Furthering analysis reveal that 60% of unstructure data is cat video.
edit: I'm sure some of theirs is from metadata, but I thought I read a while back that they were doing some graphical identification also.
Pre deep belief network I'd agree with your guess on convolutional neural networks. However, now I'd guess you'd use a deep belief network to create a network that would pick out better features than those picked out "by hand" in the convolutional neural network. (See for example [1][2])
So my money would be on some deep belief network.
[1] Hinton, G. E, Osindero, S., and Teh, Y. W. (2006). A fast learning algorithm for deep belief nets. Neural Computation, 18:1527-1554.
[2] Building high-level features using large scale unsupervised learning arXiv:1112.6209
everyone here found out about Deep Neural Networks and that is all they know.
It may take some time to match from 3 million images, but doable right? Or am I missing something here?
Fighting your machine learning algorithm against everyone else's. Best one wins the prize.
There's a lot more "interesting" competitions from different companies of course.
Unfortunately it's extinct. http://en.wikipedia.org/wiki/Thylacine
Edit: to clarify, is it safe to say that because cats look mostly alike, they would be easier to recognize consistently? or at least a good place to start?
only to people i guess. Reminds me about experiment with chimpanzees having problem to recognize faces (face recognition is a sign of intelligence according to the human thinking about intelligence) ... well, until experimenters stopped showing human faces to the chimpanzees and started to show chimpanzees faces :)
It would be interesting to see whether well-trained AI would "think" that "cats look mostly alike" compare to say dogs as it is only artifact of humans perception and of how our perception propagates into the software what we create.
This is mainly due to breeding. Dogs can vary between 8-80 lbs, depending upon breed - some will fit in handbags, others will barely fit into a car. Cats, on the other hand, have significantly less variation (between 8-25 lbs[1]).
Further complicating the problem is that we have bred some dogs specifically for facial features and shapes. An English Bulldog, for example, has a drastically different face than a Labrador.
However, the vast majority of cat breeds retain the same facial features and shape. Those that do vary (e.g, Siamese cats) differ by small amounts in comparison to dog breeds.
This is why building a "dog or cat" detector is reasonably straightforward (if it's not a cat, it must be a dog), but building a "dog, cat, or other" detector is far more complex.
[1]:http://www.petobesityprevention.com/ideal-weight-ranges/
>However, the vast majority of cat breeds retain the same facial features and shape.
i had cats for many years and with time learned to see the difference, and i'm sure that some people like judges from cat shows would see even more distinction.
It is all about model of perception (and we naturally think and talk like our, human, model is the [only] model) and how well it is trained.
>Dogs can vary between 8-80 lbs
btw, it is at least 2-180 lbs for adult dogs :)
my dog without using TV or computer (at least to my knowledge as i don't know what he is up to when we are not at home) easily recognizes other dogs of all the different breeds and sizes from the distance like across the street, etc...
And, of course, cats move in a very different way to dogs!
¹ There's argument that they're now different species, since they can no longer successfully interbreed -- for purely mechanical reasons.
Smell and sound are primary senses for dogs. Sight, not so much.
when another dog is inside a car that just stopped at the intersection?
Another issue here is that smell of different dogs is supposed to have at least some variation as well (is this feature variation bigger or smaller than variation in size?), and if one dog is downwind then another is upwind - i.e. while smell obviously plays a major role in dog's sensing of the world we just can't ascribe it all to the smell. In my experience visual recognition plays major part in many cases as well (note: i'm not arguing which dog's sense is strongest, only that there are situations when visual is basically the only one that could have brought the information)
Congratulations, you've just dramatically simplified the algorithm. This means that once you can identify a cat, you can say that any image which isn't a cat is most likely a dog.
Also, cheer up! This isn't supposed to be "useful". Who cares if it's narrow and can't be applied to anything else? It's a chance to have a bit of fun and for some people (like me) it's a chance to learn about image recognition techniques.
Oh, you mean like this picture from 1759 (you know, before TV/computers)?
http://upload.wikimedia.org/wikipedia/commons/6/6f/Louis-Mic...
Sorry, but your comment is complete nonsense. Not only about the pugs, but in fact cats vary quite a bit: http://en.wikipedia.org/wiki/Sphynx_(cat)
Oh, and the fact that some dogs are small and some cats are large.
But otherwise, I think you're on to something. Something that won't work, but it's something.