EDIT: Pun retroactively intended.
This makes me wonder. Why did people compete for the Netflix prize, [1]? Were those IP terms better?
It's not even just picking something up, it's picking one specific item out of an arbitrarily messy or overstuffed bin without dislodging or dropping anything else. Items in a bin can be wedged anywhere, in any position, in nondescript boxes only identifiable by the ASIN, or a vague natural language description. It's a hard visual logic, inference and manual dexterity problem made harder by unpredictability and harder still by the need to meet quotas, and every job has its own mess of them.
And on top of that, employees in Amazon warehouses are often cross-trained to save costs, which means you either need one robot that can unload, decant, move juice carts, pick, stow, pallete-stow, etc, at a moment's notice or else multiple dedicated robots for each task.
It's something humans do so easily that Amazon can justify paying almost nothing for the work, but still way beyond what AI and robotics can achieve. The Kiva pods seem to be the current state of the art for Amazon warehouses, and all they seem to do is move bins around (and can't even get that right sometimes.)
The problems start when things are tightly packed. One object might be in the way of your gripper so you can't take the thing you want. What happens if you drop the object? Does your robot know how to move things aside to grab the item you want?
And on top of this you have the whole issue of detecting objects in a cluttered environment: this requires state of the art semantic segmentation, robust 3D vision that will work on non-cooperative targets and a classifier that can work with occlusions and huge numbers of objects reliably.
I think it's reasonable that some items can be scanned upon entry into the warehouse, e.g. on a laser triangulation turntable that also grabs a video as it rotates. With this sort of task you want to exploit as many priors as you can find. If you look at papers (and factories), you see a lot of tricks: objects on flat surfaces that make segmentation easier, small numbers of known objects, etc.
If it could be done quickly enough with enough volume, maybe, but the constraint is that items have to be loaded off the truck, scanned and stowed into a bin (which makes them available for purchase on the site) as quickly as possible.
I wonder how they pick up books using this method?
Books are difficult because they appear to have hard surfaces on the side, which turn out to be not so hard and not suitable for gripping.
Also books, unless they're shrink wrapped, have a tendency to open which I imagine would be a pain.
Simple algorithm for that : Annoy (pun intended) towers :-)
Mimicking that mechanically is a major challenge, it has nothing to do with AI per se.
It's not a trivial challenge in any form really.
Also, there is an interesting theory that humans hands not only evolved to grab things, human hands evolved to form "efficient" fists for punching.
http://www.latimes.com/science/sciencenow/la-sci-sn-human-fi...
Reliable force feedback data at our hand's level of detail would provide a wealth of data to those making the control software and that better software could benefit from real time data. The two problems are strongly interconnected.
I would like to point that with pair of tongs you will still have force feedback. Even though is will be reduced in precision it will still be better than any artificial system I am aware.
Thing is that each of us accumulate prior data that we can use as a basis for future action, things like how much pressure is enough for grasping a raw egg without breaking it.
http://www.takktile.com/main:tech
There are other methods of figuring out how much force a hand is exerting too.
For what it's worth I for example only learned to tie my shoes when I was already 9 years old, while I had learned to read almost all by myself at 5 (my dad had taught me the letters of the alphabet).