Docker for data scientists: Introduction and use cases
unsupervisedpandas.com
unsupervisedpandas.com
But they have large down sides as well which slow us down. They are a pain to maintain as they are somewhat undocumented (you make a poc for 1 and management always wants more without improvments), a lot of edge cases cause issues which are tough to reproduce (locally sometimes impossible and waste a lot of time) and it takes them a while to start, run, etc.
This is not too tragic for nightly tests as we get the results in the morning but for tests which are started every hour, you do not want to wait that long to verify your changes work/didnt break anything. You can do these in stages, where you create different images based on the result of a previous job (run basic tests that cover base functionality that should always work, then run more in depth tests, then run performance tests at the end to ensure no significant degradation was introduced, etc..) and send out notifications asap in case of failure. The Dockerfile is essentially the documentation as you can see what is installed/configured. You can run everything locally just as it would in a k8 env. which for some reason every one always struggles with.
I am sure there are also edge cases with Docker that are a pain as well but the other selling points show it may be the right direction. You just havevto find use cases and evaluate them.
I know that the HPC clusters I've used in the past few years have all supported Singularity, but none have supported Docker (aside from our small lab cluster). Many HPC admins are (understandably) hesitant to allow non-admins access to start Docker containers (requiring root), but Singularity has no such user permission issues -- and it's faster than initializing a full VM to run a job. I don't expect that to change so long as starting a container requires root-effective permissions.
I suspect that many data scientists will be in a similar situation w.r.t HPC clusters (except for those that are using custom clouds like Seven Bridges).
Regarding HPC, from what I remember they usually have old kernels which are running (2.6) for compatibility reasons where Docker usually is not supported (unless it is backported like in RHEL).
Pachyderm does use Docker under the hood, but we don't obfuscate it away, so Data Scientists get the full power of Docker (and most of the power of Kubernetes) at their fingertips. This means you can easily grab prefabricated environments such as Jupyter or Tensorflow container images and deploy them directly. Docker is so good for packaging environments we didn't want to conceal it.
On the other hand, we felt the data orchestration capabilities of Docker were pretty lacking for Data Science use cases, so that's where we've focused our energy. Volumes are a good basic tool for getting data to your code, but they're a pretty blunt instrument. How do you split data up to parallelize over it? How do you make sure you've got the right version of data? How do you schedule new computations when data becomes available? Those are some of the use cases we solve with our distributed file system PFS.
Fighting with dependencies sucks when you aren’t intimately familiar with the ecosystem: python, node and other js, and even go.
What if a data scientist could deploy models right from browser?
Standard procedure, takes two minutes and I have a Jupyter notebook up, smartly automated base processes for dependency management, full control of environment variables and can deploy/ integrate this in any other Python setup (if I were to port IPython code to pure Python).
Is this more a thing of each his own or am I missing a crucial advantage of Anaconda?
[1] https://github.com/kennethreitz/autoenv [2] https://github.com/jazzband/pip-tools [3] https://github.com/pyupio/pyup
I've been to enough software carpentry talks where people nod and smile, and then promptly ignore the "how to be a responsible developer" advice. They don't care. They just want to run Python and plot some stuff or classify some data without worrying about multiple CLI tools. People seem to accept Anaconda because it comes across as a monolith and it's easier to understand for new programmers than virtualenvs and worrying about pinning pip installs.
That said, I don't use Anaconda on Mac or Linux.
P.S. -- Even if you are on Windows, whole SciPy ecosystem can be easily installed with pipenv, no need of conda for that now.
What benefits does pipenv have beyond storing the hashes as well? In what way is it more sane? For example, does it get around the issues with, e.g., matplotlib being installed in a virtualenv?
If you are happy with your pip and venv workflow you can happily ignore pipenv, I personally prefer it because:
- It combines all the features of pip and venv with a better UI(in my opinion).
- I can specify my environments in more detail using Pipfile [https://docs.pipenv.org/advanced/#specifying-basically-anyth...]
- It is integrated with pyenv [https://docs.pipenv.org/advanced/#automatic-python-installat...]
- The promise of deterministic builds
Oftentimes with ML and compute-intensive workloads, performance is an aspect of correctness. It is nontrivial to get the right accelerated libraries, and the right GPU libraries linking against the right Numpy version for the particular notebook you're running.
If you are doing work just for yourself, or if you have total control over the deployment environment (and your collaborator's environments), then you might be able to get away with just using pip for these things.
The Skymind Intelligence Layer helps data scientists operationalize their models on-prem and in the public cloud, and uses Docker to do that.
I would have thought most DS wouldn't be interested in learning docker just to deploy. But then again, I could be wrong. What made you write the post? Was it because you saw lots of DS wanting to deploy, or something else?
building the model and packaging it for deployment are two completely different sets of tasks, are they not?
much more useful, again in my experience, is a tool like vagrant, which allows a production environment to be provisioned onto a VM; said VM is then run from the data scientist's desktop (eg, so they can do their work with the same versions of the same tools available in prod)
From the sibling comments in this thread, it seems like there are some agree, and others don't. I was curious if there's a segment that does, and who they might be.
I also practice data science from a more engineering angle than many which further colors my opinions. Building models is fun, but I also enjoy deployment and maintenance, so I don't _want_ to hand that part off.
In contrast, I've found trying to do the same thing with Python nightmarish - simple things like having Pandas .20 on one machine and .21 on another will result in silent changes in how some metrics are calculated.
Anaconda mostly, but not entirely, solves that problem with its superior environment management (significantly better than pip IME.) But given that it's imperfect, I'll probably look more into Docker soon.
In theory, pip lets you do this with pip freeze, but I found that pip often broke on some finicky libraries, while Anaconda failed much less frequently (though still has occasional issues.) The main drawback is the same OS requirement – this makes it harder to e.g. develop on a mac and then launch a pipeline on an Ubuntu machine in AWS.
I use Docker daily for my needs. It ensures that I can easily prototype a scalable set of services on my local machine, and then deploy it to a multi-node environment on AWS, Gcloud, or our on-prem infrastructure. Good luck developing a Redis-backed ML app without Docker when it comes time for deployment. You'll get it done, but it'll be a lot slower and more difficult than if you had used Docker or containers of some sort.
Edit: Make sure to check out Yhat. https://www.yhat.com/
I set up Kubernetes to deploy Docker containers, and it took a little while to wrap my head around it. What's the painful part of doing the deploy with Docker like you've been doing?
For data scientists who write workflows as R Markdown documents and want to containerize them, you might want to check out our R package liftr: https://liftr.me/.