Neural Networks That Describe Images
cs.stanford.edu
cs.stanford.edu
"Citizen #5135: In 2015, spit on sidewalks 28x this year, jaywalked 49x, etc..."
Don't worry. Google and the CIA have your back.
/duck and covers from incoming thought police
Eh, no. Safe preventive action against a gun (or other lethal weapon, really) needs to happen before the weapon is ready for use. What if the assumed offender was wearing protection? Perfect aim wouldn't help. Lethal force might be possible to deploy - but probably not in a safe manner.
Or in the simple context of computer vision and the camera, raise warnings as individuals on screen act increasingly more suspicious. I imagine future technology will at least have intelligent security cameras that signal alarms based on the current video feed, not just when it's suddenly shut off for example.
You had to live -- did live, from habit that became instinct -- in the assumption that every sound you made was overheard, and, except in darkness, every movement scrutinized. -George Orwell, 1984.
[1] http://bavm2013.splashthat.com/
[2] video: http://www.youtube.com/watch?v=DK6KfUsVN8w
[3] slides: http://bavm2013.splashthat.com/img/events/46439/assets/a10b....
It's like those "up to XX% better" claims - "up to", not "at least" being the key phrase here.
http://cs.stanford.edu/people/karpathy/deepimagesent/devisag...
You sure? :) I definitely think the text does bias your interpretation though, I did the exact same.
> You sure?
Sounds suspiciously like something a neural net would say ...
The weird thing is how at a glance, it seems pretty much correct. And how some people here are willing to look at that and think 'automated law enforcement is clearly imminently possible'.
At this point, this software is as useful at describing photographs as a disinterested teenager who is busy trying to text. "This is a picture of my mom with some dude playing I dunno like tennis or something. Whatever."
Which is really impressive! Seriously!
But closing the gap to accuracy is really important, and it's a hard hard problem.
Nobody has been able to determine what the structure of a neural network should look like for any given problem (network type, number of nodes, layers, activation functions), how many iterations of the parameter optimization algorithm are needed to achieve "optimal" results, and how "learning" is actually stored in the network.
Statistical learning methods are obviously still useful, but I think the field is still wide open for something to emerge that is closer to true machine intelligence.
Also, 'nonbody know hows learning is stored'? You very clearly have never worked with neural nets before. Experience is stored in the form of weight values.
Where's the incorrect data stored? How can you fix it? It's in the weight values, somewhere, but you can't go and change the weight values to fix the horse/fish cascade without breaking everything else it knows.
Yes, we know 'where' the data is stored. But it's diffuse, not discrete, so we can't separate it from other data.
This is something actively being done by nn researches. And it lets us do things like take the low level audio processing part of a neural net trained on english voice data, and use it to train smarter neural nets on Portugese voice data than you couldn've without the English voice recordings.
They are reasonably successful because spammers have enough other targets that not many see it as worth the extra effort (and clock cycles) to break them, not because most of them are particularly hard to beat any more.
The average person who just wants to automate filling out your website form is still blocked, so it's not useless.
There is some recent research that suggests you can make images which are very hard for neural networks to identify, but still easy for humans.
Here is an example: http://i.imgur.com/K6AQRkV.png The digits on the right are just slightly changed to be harder for NNs to recognize.
For comparison, this is the amount of random noise needed to have the same effect as their method: http://i.imgur.com/Asnf2L8.png
Wondering whether there's any merit to sibling comments speculating this is the future of e.g. surveillance
http://www.newscientist.com/article/dn24946-google-buys-ai-f...
http://en.wikipedia.org/wiki/Knowledge_Graph
http://appft.uspto.gov/netacgi/nph-Parser?Sect1=PTO2&Sect2=H...
http://appft.uspto.gov/netacgi/nph-Parser?Sect1=PTO2&Sect2=H...
I always find that quite amusing.
Are we already at the point where NN can arrange perfect sentence when we throw bunch of words into it?
E.g.
"girl in pink dress is jumping in air."
vs
"black and white dog jumps over bar."
"woman is holding bunch of bananas." hell, I (hopefully a human) would recognize her as a male at first glance.