That was my understanding too, but seems like to handle the vast amount of image recognition required for faces and also objects which they demoed; there would need to be a large data set to learn from, and results from that learning saved.
The data set to learn from probably can be gathered from stock photos or other open source images, but can that learning be saved down to phone and used to compare against, or would it be too big?
If too big, does that mean characteristics of an image would be computed on the phone sent to the cloud compared and the results sent back?