On the other hand, you can bet that actual practice in Zalando (the authors are all from Zalando's reasearch lab) involves more regexes and retraining models on proprietary datasets and less using off-the-shelf models and hope they stick.
No one claims either that you can solve every Vision problem with a model trained on ImageNet - you'd do transfer learning, or for non-understanding problems (estimating colors and contrast or anything else that's unrelated to objects in the image) you'd use something else that doesn't involve deep learning models at all.