Suddenly, a leopard print sofa appears
rocknrollnerd.github.io
rocknrollnerd.github.io
Also, we carried out an experiment on ImageNet and the outcome was that "One human labeler (me, incidentally) with a fixed amount of training and a slightly-above average determination reached ~5% top-5 error on a subset of ImageNet test set". The media sees this and it immediately gets spun to "AI now Super-Human. And we're all going to die." It makes a lot of us cringe every time.
Many people in Computer Vision now consider ImageNet "squeezed" out of juice - we're good at texture recognition and when an object is in plain view, and we're now searching for harder tasks and more dynamic range with respect to human performance, in areas such as harder 3D/Spatial tasks, Image Captioning, Visual Q&A, etc. The hope is that these harder datasets might in turn guide us in developing models with more nuanced understanding.
It seems that this is more like the way that we learn to identify things. Then once we establish an understanding of a base class (big cat) we can apply that same model to new cats that we have never seen before with just a picture.
So I wonder to what extent you would consider this a predictable outcome from the classifier in question not being part of a subsumptive architecture --- which at a guess would look like
- glance/texture responses fed into
- boundary-recognition layers fed into
- object persistence/tracking layers
- fed into abstract scene reasoning
It seems to me, as a non-vision researcher (I mainly worked in planning and control), that the most obvious counterargument to the image being a spotted cat is based on boundary/object/scene reasoning, and that it's "reasonable" for the texture/glance layer to say "looks a lot like a cat texture".[1] https://en.wikipedia.org/?title=Subsumption_architecture
[2] I realize this may seem, superficially, anathema to deep network research, which advocates letting the network find its own intermediate levels of abstraction. But it's actually compatible in my view because Brooks advocates (again, paraphrasing quite a bit) that the separate layers should have different objective functions, and that in fact the need for different objective functions (in a prioritized order) is the cause of emergent layering in nature. "First, don't die. Second, find shelter. Third, find food etc." So one can imagine deep networks each finding their own locally useful abstractions for each objective function in the "Maslow" chain, while still having some macro architecture that tracks human-imposed design principles.
the problem is.. human vision doesn't work just by feeding a bitmap. we have structure to decode space relationships, shapes and maybe even shadow/light relations. no way we gonna see classificator working on color arrays matching our vision capabilities
However, the advantage to the texture approach is it's abstracted from a lot of other information. You don't want a classifier to say sofa, when it's a picture of a person on a sofa.
http://www.bespokesofalondon.co.uk/assets/Uploads/bespoke-so...
anyway it does work perfectly if that's what you need, but most proponent are trying to use deep nn to classify 'as good as humans do'
Context: Evolutionary algorithms and analog electronic circuits
> One thing stands out when you try playing with evolutionary systems. Evolution is _really_ good at gaming the system. Unless you are very careful at specifying all of the constraints that you care about you can end up with a solution that is very clever but not quite what you had in mind. Here power consumption is the issue. If you tried to evolve a sturdy chair you might end up with something that is 1mm tall. or maybe a fuel efficient car that exploits continental drift.
I think it's the same here: The net is never gonna better than what it needs to be, and it is probably always gonna take the easy route.
Given enough time, you'd expect the to develop clever behaviors, but instead they just fuzz-tested the sim and locked in on exploits of bugs or environment settings. They only got a bit more clever when connecting different sims running on different conditions.
Eyes already use different kinds and densities of sensors optimized for either detail and color or movement/edges. I wouldn't expect a single learning method, even after optimizing it to its limits, to be above what two or more layers of different methods could do, especially when trying to avoid exploits like the tank story.
Classic A-life! Also, not so different from the spirit of actual biology.
They only got a bit more clever when connecting different sims running on different conditions.
Diversity is very important for evolution on many levels. What many don't realize (especially, I note, evolution deniers) is that the ecosystem as a whole provides a very complex and continually varying epiphenomenal fitness function to any given organism.
However I think that's ok. Most of the fun with darwinbots is programming your own bots. They used to be (still are?) competitions where people wrote their own bots and had them compete under different conditions.
Part of the reason why a lot of these nets are trained with added noise, as well as drop-out (randomly disabling 50% of the hidden neurons, every training step).
Especially the drop-out tactic is particularly effective at preventing "exploits" of the neural net type, which otherwise appear in the form of large correlated weights (really big weights depending on other really big opposite weights to cancel out--it works, but it doesn't help learning).
Either way, adding noisy hurdles helps because exploits are usually edge cases, and noise makes them less dependable, as the region of fitness space very close to an exploitable spot, is usually not very high-ranking at all (which is why you don't want your classifiers ending up there).
And I think this is the most interesting part.
One of the most depressing things about all of the "this image recognition algorithm performs better than humans on this task" is the idea that we've pretty much solved the problem, and it's just a matter of some more optimization and tweaking to handle a few edge cases.
This kind of problem, where the dominant solution simply gets it so wrong, and the problem cases are uncommon enough that any statistical solution is generally going to treat them as noise, reveals that in fact that there is likely plenty of room for entirely new, novel ways of approaching the problem to handle these kinds of cases better.
It's actually more exciting that there's so much more to be done, than to say "well, it's basically a solved problem, we just need to do some tweaking and optimization."
Indeed, discovering these "broken" edge cases is exactly what we need to converge upon a more correct solution.
This is a classic story of a neural net failure.
The net was able to find tanks hiding in the trees with amazing accuracy. Too amazing. It turned out the photos of the hidden tanks were all photographed on a cloudy day. The images without tanks in a clear day.
My nitpick:
> When each student was given a heavy book of MNIST database, hundreds of pages filled with endless hand-written digit series, 60000 total, written in different styles, bold or italic, distinctly or sketchy. > ... > So, are you going to say that was not the case?
I understand the point the author is making. Human brains are really good at taking limited examples and correctly extrapolating them to new cases. That is, of course, the goal of intelligence. Machine Learning has gotten better at this generalization, but has a long way to go. And ConvNets as they exist today will not achieve that, no matter how much training you perform on them.
This specific example is inaccurate though. Let us aggressively simplify and low-ball by saying that humans see at 24fps. Humans of course don't see in discrete frames, but this simplification doesn't detract from my argument and makes quantifying easier. So, if you give a human a single page of numbers, and they look at it for an hour, they have now seen >86k examples. That's 86k examples with twitching saccades, and from both eyes. That's in just an hour of looking at numbers.
Prior to being given that page of numbers, most children will have been alive for 4-5 years. That's 3 billion examples from a wide variety of subjects (we ignore sleeping cycles, because we're already low-balling this fps figure, and because the brain is still learning and visualizing during sleep).
And humans are born with a pre-built visual cortex. Edge detection, gradient detection, etc. are all already built for us. CNNs learn that from scratch.
The author's real point is still valid, though, don't get me wrong. I'm just nitpicking.
You can show me novel symbols, with me only looking at a single example of each for a few seconds, and I can manage good categorization.
Test your humanness; draw these symbols:
"Like an E but rotated so the prongs point upwards"
"Like a snake but with two heads. Snakes down, up, down, up, down."
"Like a walking stick with the handle pointing left and looping back around."
(answer for A: Russian letter Sha) (answer for B: Kannada letter Uu) (answer for C: Tamil vowel sign I)
And then it was able to correctly recognize 7's and 8's, despite never having actually seen one. I'm simplifying somewhat, but it was super cool.
I don't know why people are so focused on one-shot learning, or think that NNs can't do it. Neural networks learn features from lots of (possibly unlabelled) data. That's the whole point. Once you have those features, you can use them for all sorts of things. You can show it an image, and then measure how close other images are too it. Thereby learning from a single example.
In the words of Joshua Tenenbaum and coauthors, "human children learning names for object concepts routinely make strong generalizations from just a few examples".
You can check this out for yourself on the brilliant illustration that went with it: http://i.imgur.com/5axtXSo.png From Tenenbaum J.B. et al, "How to grow a mind: statistics, structure, and abstraction," March 2011, Science, DOI:10.1126/science.1192788.
I myself have researched leopard spots since I painted our toilet floor in them. It's a lead sheet, and the paint had worn off, which probably wasn't the healthiest thing. My housemates had filled the toilet with memorabilia from an African trip, so leopard-print paintjob it was.
Which entailed looking up leopardprint online. Very little of which actually looks like leopard rosettes, and now I have a problem with almost anything trying to pass itself off as leopardprint. Anyway, I can't say that my paintjob is a particularly good reproduction, but at least it's 'spiritually correct'... :)
My son keeps telling me that infants are fine with, say, a truck transforming into a clown (when it emerges from the other side of a visual barrier) but not with it transforming into TWO of something. Apparently, babies subjectively experience this (visual transformation) all the time -- mom moves a plate and what seemed like a big circle is now a flat line or whatever.
So humans apparently get tons and tons of experience with visually mapping 3d reality to mere 2d imagery. I have been thinking somewhat about this of late, in terms of physical attractiveness or "image" -- that pictures of a woman posted on a blog capture a 2d version of her but people interacting with her are interacting with a 3d living, moving creature who also has smell and a voice and her movements may be elegant or may be not elegant. Which is a thought process relevant to a project of mine, something people here surely will have no interest in. But where it is relevant to this article is that we are doing this wrong: Humans have thousands of hours of practice of looking at 3d reality and figuring out how it to interpret 2d images as representative of that 3d reality. Image recognition software is just dealing with 2d images. I don't see how it can hope to compete. Humans don't come preinstalled with the software to make that distinction. We acquire it with enormous repetition.
When do we make a robot and give it some baseline parameters and a learning algorithm (and set it loose in 3d reality and have to learn)? That is when we can get scared about human like AI that can compete on image recognition.
I suppose that's why Banach–Tarski is considered a paradox.
Edit: There's a comment about invariance in this thread [1] and apparently CNNs are not invariant under rotation.
Giving it an image that we know has all the relevant details of a sofa, but it likely won't have close matches in the dataset, can give us an idea of how clever it is.
https://upload.wikimedia.org/wikipedia/commons/7/77/Martian_...
We're doing humans wrong. Maybe not all wrong, and of course, humans are extremely useful things, but think about it: sometimes it almost looks like we're already there. There always going to be an anomaly; lots of them, actually, considering all the things shaded in different patterns. Something have to change.
I agree that we aren't there, but we'll never be there, every system can be fooled, its just a question of 95%, 99% or 99.99%
The fact that all classifiers -- including human beings -- fail in some cases is a separate issue. The goal is to create a computer classifier that succeeds and fails in the same cases humans do.
I don't follow. If you asked a human what the linked image looked like, they'd likely say a face, but if you then asked them what it actually was, they're all going to change their answer to a rock, even specifically a rock on Mars (if given a colour version of this image).
It's true that humans see patterns that aren't there, but does that detract from our ability to recognise objects?
I tried the unrotated sofa image on Wolfram's ImageIdentify and it correctly identified a settee [1]. So it presumably gathered that from the shape of the image rather than the pattern. It is peculiar though that it can't see the shape under a simple rotation. Or perhaps the margin of confidence levels between sofa and leopard were so narrow that a rotation was enough to tip it in favour of the leopard? I'd be interested to see the inner workings of this.
I kept trying different ones and it kept identifying as "Bicycle Saddle"...
Scary and fascinating.
> When Vladimir Vapnik teaches his computers to recognize handwriting, he does something similar. While there’s no whispering involved, Vapnik does harness the power of “privileged information.” Passed from student to teacher, parent to child, or colleague to colleague, privileged information encodes knowledge derived from experience. That is what Vapnik was after when he asked Natalia Pavlovich, a professor of Russian poetry, to write poems describing the numbers 5 and 8, for consumption by his learning algorithms. The result sounded like nothing any programmer would write. One of her poems on the number 5 read,
> He is running. He is flying. He is looking ahead. He is swift. He is throwing a spear ahead. He is dangerous. It is slanted to the right. Good snaked-ness. The snake is attacking. It is going to jump and bite. It is free and absolutely open to anything. It shows itself, no kidding. Brown_Cornerart
> All told, Pavlovich wrote 100 such poems, each on a different example of a handwritten 5 or 8, as shown in the figure to the right. Some had excellent penmanship, others were squiggles. One 5 was, “a regular nice creature. Strong, optimistic and good,” while another seemed “ready to rush forward and attack somebody.” Pavlovich then graded each of the 5s and 8s on 21 different attributes derived from her poems. For example, one handwritten example could have an ‘‘aggressiveness” rating of 2 out of 2, while another could show “stability” to a strength of 2 out of 3.
> So instructed, Vapnik’s computer was able to recognize handwritten numbers with far less training than is conventionally required. A learning process that might have required 100,000 samples might now require only 300. The speedup was also independent of the style of the poetry used. When Pavlovich wrote a second set of poems based on Ying-Yang opposites, it worked about equally well. Vapnik is not even certain the teacher has to be right—though consistency seems to count.
http://nautil.us/issue/6/secret-codes/teaching-me-softly
That article in turn reminded me strongly of "Metaphors We Live By" by Lakoff & Johnson, and the works they have written since, where they claim that humans make sense of the world using systems of rich, conceptual metaphors. As I understand, the work is well-known to machine learning researchers.
I'd just like to answer the recurring objection: yes, our visual experience contains a lot of frames and that seemingly refutes my MNIST example; however, you do forget about the other part of a supervised dataset, namely labels. Do we have a label provided to each thing we see in our life? Obviously not. How much time do you need to familiarize yourself with a new entity, like an unknown glyph or symbol? Can't provide a concrete example, but I guess a single math class was enough for all of you to recognize all the digits the next day. You can test it right now by looking into some unknown alphabet and then looking into it again upside down - you'll recognize it perfectly, except for mental rotation issues (which occuur even for well-known letters and symbols).
Also, being able to consciously recognize letters is relatively easy, but the normal reading process, with which people recognize well-known letters and instantly unconsciously convert them to sounds, does require quite a bit of repetition of those letters before it starts to kick in...
I'm curious what makes you think that. My experience with what's going on at my sons school is telling me that the children spends a massive amount of time on getting recognition of digits and letters right.
I think you look at your subjects wrong - don't pretend your computer is an adult (which had learned most of its life) - rather consider him an infant learning letters/digits/objects for the first time. Doing so you might come across a similar learning curve to the one you have described. With the additional case we (or at least I) don't know how to bring the computer to the level of a fully grown man.
As for the problem presented in CNN, if the problem is not having the structure, why not gray-scale the structure as a secondary level for the CNN?
I'm not really from the field so excuse me if this was complete BS
This image: https://i.imgur.com/2aCqMx2.png
And here are the results: https://imgur.com/a/8ndyq
This doesn’t really prove anything, but I thought it was interesting. It is of course, unreasonable to expect ML algorithms to perform decently, so far outside of the space they were trained on.
But I suspect that part of the reason they don’t do well is that they are purely feed forward. Humans also don’t see the image at first. It takes time to find the pattern, and then everything clicks into place and you can’t unsee it.
This might have something to do with recurrency. But more importantly, information feeds down the hierarchy as well as up. Features above, give information back down to features below. So once you see the dog, that tells the lower level features that they are seeing legs and heads, which says they are seeing outlines of more basic 3 dimensional shapes, and so on.
I think it also requires a descent understanding of 3d space, to fit the observed pattern to 3d models which could have produced it. I’m not certain if regular NNs observing static images, are optimal for learning that.
More here: https://www.reddit.com/r/MachineLearning/comments/399ooe/tes...
There are no sofa's in this list, the closest thing I can find is a "studio couch, day bed": http://imagenet.stanford.edu/synset?wnid=n04344873
Best guess for this image: cat bed furniture
https://goo.gl/vXwSajBut high-level notions like a Jaguar is a cat-like animal aren't necessary to perform well on an N-way classification task like ImageNet.
What's more important to note is everybody knows there's plenty wrong with a pure appearance-based approach like CNNs. Every few years a new approach pops up that is based on ontologies, an approach inspired by Plato, etc, but these systems require a lot of time and effort. More importantly, they don't perform as well on large-scale benchmarks. In the publish-or-perish world, you can jump on the CNN bandwagon or start reading Aristotle's metaphysics and never earn your PhD.
If it doesn't work at all, or isn't a new idea, that's different.
I would propose that for this leopard problem, instead of just skewing the images, you also performed transformations on the COLOR and put the images back into the training set.
Maybe applying certain filters, such asdimming the saturation or contrast of images, so that the contrast between the leopoard spots were less visible (i.e. "A Leopoard in low lighting") - maybe this would force the neural net to learn more than just its print.
Knowing the right set of color filters to apply to all images could be tricky though.
(I've seen the comments like https://news.ycombinator.com/item?id=9584325 and watched the lectures and youtube walkthroughs, but they're all theoretical and I'm looking for documented code to go along with that theory)
The net correctly identified "leopard". Was it taught about sofas? Who knows, maybe Sofa had a high score as well on the output.
Or, look at the Dalmatian/Cherry picture. The net identified "Dalmatian" which is a 100% valid response! But whoever labeled it wanted "cherry". The picture is 50% cherry 50% dalmatian.
Pictures often have more than one element and a pure ConvNet is "one picture to one label"
If part of the network is trained on the concept of a cat, and whether or not an image is a cat is fed into training of the leopard, it seems like the problems would be avoided. Or is the notion that with enough training data and deep enough networks the concept of "leopard is cat" will be learned?
The problem with attempting to understand intelligence by reverse engineering the human brain is that we cannot know a priori which aspects of the human brain are necessary for intelligence to arise, and which are merely consequences/side effects of biology and chemistry. Once we discover some technique that works in a practical setting (e.g. on ImageNet), then it is fairly straightforward to find the biological analogy in the brain.
In fact, Geoff Hinton explicitly advocates an approach of "try things, keep what works, and figure out how it relates to the brain". The inverse is like finding a needle in a haystack.
There is also clearly no comprehension of the importance of the topological (circuitry) defining a neural network. We always assume a fully connected network, and draw the out as such, but we don't stop to consider that many of those Wijk interactions are completely spurious, meaning they have no information bearing role.
The purpose of training a deep neural network from data is to automatically discover what the topological circuitry of the network should be, rather than engineering it by hand. In the brain, some prior knowledge is encoded via genetics, while the rest is learned. The effect of sparsity of the weights in deep neural networks is an active area of research [1].
If you strip them away you'll start to reveal the underlying circuit at work. I've published theoretical results using artificial gene networks, but the results should be similar for ANNs.
Very interesting. If I understand correctly, the cost you are attempting to minimize is phenotypic variation, which you measure as the gross cost of perturbation (GCP). Would this cost be analogous to sensitivity to adversarial examples in the case of convolutional neural networks [2]?
Regardless of how computationally expensive NNs may be now, wait a few years, and then train millions of them on different classes of objects and run them concurrently to identify new pictures.
Also, the author spends the first section of the article determining that it is in fact a jaguar-print sofa (which the model also confirms) but continues to throw around the word "leopard". They're not making it any easier for the future machine learning algorithms that try to identify an image by the text surrounding it. ;)
[0] http://www.tineye.com/search/4c4ce7b6558e8d3c4dd443439e80556...
If one paired the present classifiers with Amazon Mechanical Turk, just providing one bit of information -- "is-it-an-animal?" -- I wonder how well the current classifiers would fare in relation to human beings?
(1) - Ironically, the more "cosmopolitan" people become, the quicker they are to jump to such conclusions!
This, given the leopard couch, returns "brown leopard print couch".
The models have room for improvement, but it's not clear to me that larger datasets won't solve the problem. Larger datasets and more processing power is exactly why neural nets have surged in effectiveness recently. Who knows how much further current models can go with more data and processing power?
This sort of works on simple situations without too much edge noise. It's been used for industrial robot vision, where what matters are the outside edges of the part. It's not too useful when there's clutter, occlusion, or noisy textures.
More recent thinking is to find surfaces, rather than edges. This works well if you have a 3D imager, such as a Kinect. You can get a 3D model of the scene. Occlusion remains a problem, but texture noise doesn't hurt.
[1] http://homepages.inf.ed.ac.uk/rbf/CVonline/LOCAL_COPIES/GOME...
It would seem CNNs were a significant step up, but the author hints at inferring structure as the next tack to take.
We certainly see the same thing with autonomous vehicles. Given very accurate mapping and a particular set of environmental and type-of-road conditions, cars can do so well that it's tempting to say they're 95% of the way to fully-autonomous. But dump them in a Boston snowstorm and you see they're really not even close. (Which isn't to say that bounded use cases can't be very useful.)
But it might be different this time...
Right now we're at the stage of the "seven blind men and the elephant". Over time, our eyes will open and things will start making sense.
(A lot probably has to do with switching to more data-based approaches.)
If not, why else would this exist http://www.nvidia.com/object/tesla-supercomputing-solutions....?