1) I'm building this primarily for a robot, so my impression is that by building up a set of primitive features, and building up on top of them, will be more efficient for soft real-time use cases.
2) I have done what you suggest (more or less), for example as a test with the MNIST Digit Classification data. I was only getting approx. 95% success rate using this basic technique, which at first glance appears impressive until you realize what is considered "state of the art".
3) My goal is to also incorporate additional layers of data (which I'm under the impression will help improve my overall success rate), such as edge detection, corner detection, and most importantly 3d data (ie. distance from camera based on data extracted from Google Tango), so if I'm correct about #1 where this will be a more efficient calculation, incorporating all the additional data (and extracting features from that as well) will slow things down only incrementally.