Evaluating recommendation systems is hard because you actually require a human in the loop. Even worse, giving the recommendations alters the human behavior. Then you need to think what metric are you going to use. For training you will most probably use a proxy metric that correlates. Maybe you want to optimize different metrics and they actually need to be balanced. Then there are lot of confounding variables: maybe a better UX will improve the metrics than a better algorithm, or a change of products.
It seems that with big enough data you can improve old models with deep learning but I think recommenders are very far from similar gains to other fields (NLP or CV for example). And most companies don't have that much data.