I saw a demo of a deep convnet (I assume) running on a tablet at a conference recently. While it's limited to the vocabulary that the net is trained on, it's seriously impressive seeing this stuff work in realtime.
One of the other speakers was giving a tech demo using his webcam. He was looking around for a mug to demonstrate that the classification was good and could work quickly. In the meantime the camera was looking behind him on stage and correctly classified the image as "theatre curtains". It was particularly cool because image processing results are often cherrypicked to show optimal performance and you learn to be skeptical.