From the sibling comments in this thread, it seems like there are some agree, and others don't. I was curious if there's a segment that does, and who they might be.
I also practice data science from a more engineering angle than many which further colors my opinions. Building models is fun, but I also enjoy deployment and maintenance, so I don't _want_ to hand that part off.
In contrast, I've found trying to do the same thing with Python nightmarish - simple things like having Pandas .20 on one machine and .21 on another will result in silent changes in how some metrics are calculated.
Anaconda mostly, but not entirely, solves that problem with its superior environment management (significantly better than pip IME.) But given that it's imperfect, I'll probably look more into Docker soon.
In theory, pip lets you do this with pip freeze, but I found that pip often broke on some finicky libraries, while Anaconda failed much less frequently (though still has occasional issues.) The main drawback is the same OS requirement – this makes it harder to e.g. develop on a mac and then launch a pipeline on an Ubuntu machine in AWS.
I use Docker daily for my needs. It ensures that I can easily prototype a scalable set of services on my local machine, and then deploy it to a multi-node environment on AWS, Gcloud, or our on-prem infrastructure. Good luck developing a Redis-backed ML app without Docker when it comes time for deployment. You'll get it done, but it'll be a lot slower and more difficult than if you had used Docker or containers of some sort.
Edit: Make sure to check out Yhat. https://www.yhat.com/
I set up Kubernetes to deploy Docker containers, and it took a little while to wrap my head around it. What's the painful part of doing the deploy with Docker like you've been doing?
I would have thought most DS wouldn't be interested in learning docker just to deploy. But then again, I could be wrong. What made you write the post? Was it because you saw lots of DS wanting to deploy, or something else?
building the model and packaging it for deployment are two completely different sets of tasks, are they not?
much more useful, again in my experience, is a tool like vagrant, which allows a production environment to be provisioned onto a VM; said VM is then run from the data scientist's desktop (eg, so they can do their work with the same versions of the same tools available in prod)