Machine Learning Containers Are Bloated and Vulnerable
deep.ai
deep.ai
Software expands to occupy all hardware available. If you throw more hardware at the problem, you get bloated software.
It happens everywhere. Powerful clients led to bloated web apps. Huge data centers led to "micro service and devops" bloat. How many layers of software do I need to run code that actually handles business logic? 'Everything-as-code' sounds great when your tool box is knowing how to write code.
95% of the actual, real world problems could be solved by spreadsheets and email. But when you pump productivity gains into capital returns and keep people working the same hours as a century before, busywork goes through the roof.
Bloat everywhere.
"Software expands to occupy all hardware available. If you throw more hardware at the problem, you get bloated software."
\/s
For example, the very popular oobabooga /text-generation-webui repo on GitHub has a Dockerfile that simply won't build.
Oh, maybe it was possible to build it once at some point in time, but then some Python package got updated, and it's game over.
It's astonishing how fragile the reproducibility of everything in the Python and ML world is. Even with the help of Docker, things break within weeks, and then stay broken.
1) didn't link the files up to the root - (in the docs - didn't read)
2) needed to specify my CUDA version in .env
3) using docker on WSL and had a hard memory limit set so the build got oomkilled
Now it's working fine
> jupyter-repo2docker is a tool to build, run, and push Docker images from source code repositories.
> repo2docker fetches a repository (from GitHub, GitLab, Zenodo, Figshare, Dataverse installations, a Git repository or a local directory) and builds a container image in which the code can be executed. The image build process is based on the configuration files found in the repository.
> repo2docker can be used to explore a repository locally by building and executing the constructed image of the repository, or as a means of building images that are pushed to a Docker registry.
> repo2docker is the tool used by BinderHub to build images on demand
There are maintenance advantages and longer time to patch with kitchen-sink ML containers like kaggle/docker-python because it takes work to entirely bump all of the versions in the requirements specification files (and run integration tests to make sure code still runs after upgrading everything or one thing at a time).
What's best practice for including a sizeable dataset in a container (that's been recently re-) built with repo2docker?
Most containers are built with layers and removing from one layer doesn't make things smaller.
older versions of docker don't help with squashing (except with experimental turned on)
I think you might be able to do with multi-stage builds and COPY
but... it's not valued