1) Once the training is done, the prediction part is simply matrix multiplications - I would think that's probably faster and far easier to parallelize than any other advance feature extraction algorithm you will be performing real time on the image.
2) Did you try using CNN?
I might be grossly wrong here, but I have gone through this path before, but I have come to the conclusion that clumsy hand-crafted feature engineering on images usually performs worse than letting the algorithm figure out. It's much harder to scale and you will end-up coming across a lot of edge cases.