Run More Stuff in Docker
jonathan.bergknoff.com
jonathan.bergknoff.com
Docker is for distribution of applications when deploying them to servers. As a developer, it's amazing at that and have brought peace and joy in devops. Let's leave it there, shall we?
Help me understand the WTF here, and also how that’s meaningfully different from an OS X app?
Have you been running Docker on your system? Have you tried running: `docker system prune -a`? It's gonna print something like Total reclaimed space: 31.2GB.
Maybe a better example is the Tensorflow image, which is close to 4GB for the GPU version. It's also the kind of thing that would be a pain to rebuild from scratch if you wanted to find a way to save space.
> the Tensorflow image is really big
This is kind of meaningless without comparing it to how large Tensorflow is with related dependencies if you just install it centrally?
I like to think I am somewhat proficient in using Docker correctly by now, but I am still discovering new tidbits of practical knowledge, tips and good practices every few weeks. And I always want more. :)
Meanwhile Docker adds many more restrictions and limitations and is not a good fit for consumer-focused interactive GUI applications which are not the web and console apps that work best with Docker.
Sure you could do the same with chroot manually and there are many other tools that does the same thing, but somehow the docker way of representing these concepts are more graspable for average user, and much more widespread use.
- Dedicated RAM for all your docker apps (which you have to partition manually)
- Slow startup time for the first docker app you run each time you reboot, as it boots the linux VM.
- All syscalls run through the VM's emulation layer, which is way slower than native
- No access to the system's native UI toolkits. (Docker apps can't dynamically link to cocoa for obvious reasons.) We can probably work around this by wasting a huge amount of developer time making a networked bridge or something silly but ??
- No easy way to load and save files (the VM doesn't have filesystem access)
- All the other downsides of docker - like it gobbling up all your disk space with images that are no longer used
- Lower battery life, because the host can't easily sleep the VM's scheduler. The CPU can't enter low power state when nothing's running.
Please don't do this. We already have a special kind of container for "self contained program on disk". Its called a statically linked executable. They work great. They're fast, small, have access to everything the host system provides. You can SHA them if you want and you can easily host them yourself using a static web server. My computer's responsiveness is more important to me than your shiny docker shaped toy.
When I say stuff like this in job interviews, I get eye rolls and don't end up getting a job.
The industry is so far up its ass lately, no one even dare stating the facts in public. /rant
The next time I have to copy/paste the same utility to a new microservice because [insert corporate reason] I’m going to scream.
As a former sysadmin, the less actual admin I need to do the better.
Wait, you’re self hosting multiple apps on one system, and not using K8?!? Your apps can’t autoscale or do blue/green upgrades or any of the cool stuff you are now mandated to do by the cargo cult. How do you sleep at night??
/s
This is ultimately what drives me insane about the tech sector: so many choices are trend based and cult like when any sort of technical discipline shouldn't be.
My website used to be in an S3 bucket, I was unhappy with it.
My personal website is low-traffic, it’s not going to be very hot in any CDN caches. Based on my experiences with S3 performance, adding another layer of cache misses in front of it is probably just going to slow things down.
The suggestion was more for the HTTPS part and not the performance part of your complaint. It's simple enough to set up, pay-as-you-go, only requires one vendor, and should be very scalable.
I'd just as soon throw up nginx on a Linode, myself.
When I stare at docker long enough I wind up at "why couldn't this be a static binary" or "this would be easier to secure if it was it's own VM".
A static binary is much better than a Docker container, I agree. Unfortunately a lot of useful apps can't or won't ship as a static binary (Java, Python, PHP apps), and Docker makes them behave much more like a static binary.
Same applies to Python bundlers like py2exe.
like http://programming-motherfucker.com but against triple-buffered-virtualization.
This, and the other VM-related criticism, is true, but only on Mac & Windows. Running a Docker container on a native Linux system carries very little overhead.
At the end of the day if you have a precompiled self-contained artifact Docker is just a convoluted way of calling exec.
You can't really reproduce that with any popular desktop operating system. Even if you could the interesting thing about docker is that starts with a default deny environment built on the principle of least privilege.
A few comments on how insecure the typical Docker recommendations are:
> Sandboxed - security claims about Docker have always been controversial. Simple Unix/BSD constructs like chroot/jails are far simpler and they are reliable
https://news.ycombinator.com/item?id=25547629
> However, I don’t agree with using it for sandboxing for security. Especially if you are giving it access to the X11 socket, as the author does in the examples. It might not be obvious, but the app will have access to all your other open windows.
https://news.ycombinator.com/item?id=25547551
> Fun fact, most docker hosts will allow access to all your files anyway! (specially true on docker for mac, which all the cool kids(tm) here are using). Even if you restrict container host-FS access to a source repo dir, mind rogue code changing your .git hook scripts in there or you might run code outside of the container when committing ;) > Another slightly relevant fun fact, USB is a bus. That means that any device can listen in on any other device. And USB access is given by default to some X-enabled docker (--tty something), and to most virtualbox machines (including the hidden one running the fake docker linux host on docker-for-mac), and more recently Google-Chrome. ;)
If you’re running Docker, you’re running Linux, and you can use the same kernel features as Docker to sandbox applications yourself. This way, you only incur the overhead of a separate network namespace and filesystem if you actually need it. Services can be sandboxed with a few lines of configuration in a drop-in unit, and applications can be run with a shell script that calls `systemd-run`, just like the `docker run` shell scripts suggested in the article.
As slow as it is some people apparently enjoy running docker from win / mac
Linux desktop apps running in docker would work better because there's no VM in between the native environment and the application. You'd still need to forward GUI calls across the LXD container boundary. That used to be easy with X11 forwarding, but AFAIK wayland removed that feature. I don't know enough to guess how difficult that would be to implement.
But, if you solved that it should be mostly smooth sailing. Linux desktop apps distributed via docker would just be the same apps, but unnecessarily bigger, with probably less access to the filesystem. They would be harder to patch for security errata. And they would leave extra junk in your docker images cache. Docker's sandboxing would be nice though.
Doesn't this make OpenGL, Vulkan, and parts of glibc break?
For many reasons, none of that is available via docker.
The GP is also referencing the fact that even if you somehow had access to cocoa from your linux binary inside your docker VM, the docker image doesn't / shouldn't have access to your host (macos's) filesystem. So the save / load dialog wouldn't work properly anyway.
I'd expect someone building many tools like the author of the original piece is using the same shared image as the base for most of them?
- They're proprietary. Copying them from macos onto docker hub violates the macos software license.
- They change with each version of MacOS. You can't mix & match them.
I guess you could mount them into docker's filesystem, but then whats the point of using docker? And even if you did that:
- They're macos mach executable files. Linux (and therefore docker) doesn't know how to run mach executable files.
- Even if you could somehow embed them and get them to run, the libraries wouldn't work because they expect to be making syscalls to Darwin. They can't do that from inside a linux virtual machine.
You could probably make a weird RPC proxy involving a native macos process receiving network commands. But it should take a herculean effort to make it work at all, and even if you got it working it would probably be buggy (since everything would suddenly become async) and slow.
If you are running a Docker image on MacOS or Windows, you first have to start docker which is itself a Linux Virtual machine.
Docker is a great dev tool, but if you don't need it, it's a ton of bloat.
A macos app is running in a sandbox and runs in a conceptually similar way to docker.
Go look in ~/Library/Containers
also look at the filesystem under <appname>.app
MacOS apps contain all the app. Mac Apps leverage all of the OS functionality they can. They are strongly tied to MacOS and rely on it for most functionality. When I upgraded to Big Sur, the Mac Apps that run on my computer adapt to the changed OS libraries and often present differently.
Docker apps contain their own complete environment. They are deliberately engineered to disassociate from the base OS.
Electron is a sort of middle ground largely ignoring many system libraries, but using others.
The only thing Docker has in Common with Mac Apps is the fact that they keep associated files bundled together.
Should I respond metaphorically?
> A macos app is running in a sandbox and runs in a conceptually similar way to docker.
A Mac App fundamentally has access to system libraries and leverages those. A Docker container is designed to ignore the system and builds its own environment.
If you run a Docker container on a Mac or Windows, you are now running 3 operating systems. The host OS, the VM, and the Docker image.
This is not the same as a Mac App. Literally, figuratively, or hypothetically.
Minimal "distrofull" images are in the 50mb range. But often you don't really need a real distribution inside your docker image.
See https://github.com/GoogleContainerTools/distroless. But distroless is not about one specific tool or base image, it's a paradigm that addresses precisely what you say (without throwing away the whale with the bathwater)
I'm not stating that I believe it is a good idea to run desktop apps in Docker containers. It is not a good idea. But it is also not true that if someone would do it, it would necessary lead to a bloated filesystem.
Docker will use a lot of disk space when using multiple base images.
By the way, while I think you should run more stuff in Docker, I also think it is still totally reasonable to have your "main language stuff" installed "normally" (not in Docker). If I'm a Java developer, I'm "happy" to deal with Java's dependency b.s. on my normal system. Meanwhile, if I ever have to touch Ruby, I do not want to deal with Ruby's dependency b.s. on my normal system (rather run that in Docker). And vice versa, as I'm sure.
Oh yes, it would make me "happy", too.
Only if they’ve all chosen the same base image.
And if application developers could all agree on a fixed base image with fixed versions of dependencies, that’d be a Linux distro and we wouldn’t need docker to begin with :)
I think the beauty of docker - and this should spread elsewhere - is that everything starts with one text file.
The crufty part is all the crap that has to be added to the docker run commandline.
That! It is codified. If people had their main machine codified, Docker would maybe be less of a benefit. And I'm sure a lot of people here have that. But a lot of my colleagues don't. So, I give them a Docker image... instead of explaining the same Java developer for the umptied time how to make a virtual environment for my Python program.
Also, docker doesn’t solve the DLL hell problem, just pushes it out of sight. The nix package manager is something that actually solves it and it should be promoted and leave docker to things it is good at - containerization.
I think the necessity of a VM when using Docker on Mac and Windows is the primary reason that running your “normal” apps in a container isn’t the right move.
Again, not advocating Chrome in a container, I don’t even run Chrome outside of a container. I just think it’s odd to get hung up on these sorts of resource requirements given the state of computing.
Related: Chrome on 4GB of RAM sounds painful. Thoughts and prayers to those folks.
However a nice SSD with a respectable TBW value is not still cheap. a 860 Pro is almost twice the cost of a 860 Evo. Pro provides twice the TBW value.
> My internet connection also makes downloading a large docker image no bigger of a deal than downloading Chrome, YMMV.
Not everyone of us has pipes that fat to our homes which provide sub 10ms pings and almost LAN-speed access to rest of the world. My office workstation's network is limited by my network card but, my home has a much slower connection.
I wish that internet on this planet to be a full-fat-tree network but, we're not there yet.
TBW is almost never a concern for desktop users. The Evo has 600 TBW endurance per TB of storage, that would be a full disk rewrite every day for two years.
You will never download enough Docker images for personal use to burn out your 860 Evo before you would have replaced it anyway. (For those unfamiliar with Docker, spinning up a 500MB image ten times doesn't write 5GB to disk!)
You're right however, most of the people who'll use this kind of setup is not ordinary desktop users.
My desktop has 4 disks (2 SSDs and 2 HDDs). My write rate for "Home" SSD is 3TB/yr. To keep the value low, I've moved VMs, big downloads and other stuff to one of the HDDs. System is on another SSD and it's cumulative write was about 3TB in 8 years but, I moved logs and high-write portions to another HDD to keep that value low.
3-4TB / year on the other hand is pretty in line with a Windows 10 installation's behavior when used by a normal desktop user, as intended.
>You will never download enough Docker images for personal use to burn out your 860 Evo before you would have replaced it anyway.
Considering other stuff I do, I could easily double or triple the amount of writes in my Home SSD, but VMs and other stuff already can sit on the RAM once running so there's no speed problem.
While write amplification is not a big concern anymore, building software and other small file operations can accumulate fast, so I still can't trust a SSD blindly.
On the replacing of drives, while a dd or rsync is pretty straightforward for a seasoned Linux user, I don't prefer to change hardware just for the sake of it, or abuse it because it's cheap and can be replaced on a whim. At the end of the day, I'd rather use my system efficiently rather than recklessly both in terms of resources and endurance, because being able to rely on your system is underrated imho.
Because of my Job, we torture systems up to and beyond their design limits and a little optimization can go a long way in these scenarios. I like to apply that knowledge to my systems to extend their useful life.
TBW is 600 TB, but you do 3TB/year across four disks, and I'll be generous and say you do .75/SSD/year.
so in 750 years you'll hit the rated TBW?
If you did a full 3TB/year on the one SSD, you hit it in 200 years?
Unfortunately, neither of my disks are that big. Home is on a 256GB SSD which boils down to 300TBW. System SSD is a bit older 120GB OCz Vertex 3. This model doesn't have a TBW rating.
As a result, if I don't do anything heavy, it'd last for a century in the best case. I'm not sure about OCz though. It reports 100% life remaining but it's from the skunkworks era of the SSDs so, I can't be sure for anything.
The numbers climb very fast when you start to develop stuff and enter compile -> test -> debug cycle. So, I'd rather have that endurance and use it while developing software rather than eating it while doing daily stuff.
Anandtech was the best site for that era with their testing.
It has no temperature sensor so it cannot compensate for temperature or understand its environment. It doesn't have a TBW rating (IIRC) and reports writes and reads in GiB. So it's somewhat limited when compared to today's drives.
OTOH, it's pretty dependable and stable so far.
[0]: https://www.anandtech.com/show/4256/the-ocz-vertex-3-review-...
Also the engine runs as root and takes commands from normal users which I always thought was a no-no but I guess docker is ‘special’ - and should be able to do whatever it likes. Containerisation is a fine idea and concept - but I think there are still big caveats.
I’ve been using podman a lot and I hope it becomes more commonplace. I hear that there’s a lot of SV politics and drama going around surrounding the various companies and backers which I have exactly 0 interest in though...
With Flakes [2] (experimental feature), you get full reproducibility.
The documentation is spotty and there is a considerable learning curve, but I've switched to NixOS on my laptop and desktop early this year and am mostly very happy with it.
That doesn't cover sandboxing though. I would actually agree that sandboxing / restricting applications (Mac OS style) would be sorely needed on Linux desktop.
Flatpak and Snap are trying to do this. A much saner approach than trying to find or maintain up to date, trustworthy Docker files for applications.
You might be interested in this fascinating exploration of attempting to bend Docker into better caching and composition by injecting blocks of Nix packages as individual layers:
The thing the upstream Linux distributions are missing are a lockfile with the hashes of installed packages. Programming languages figured this out (go.sum, package-lock.json, etc.) but distributions have not. Thus, people are often running "whatever" in production, because they simply don't have the ability to lock dependencies properly.
I assume Nix solves this problem, and people should pay attention to how important it actually is.
Are solutions like source2image, jib, or the nix dockerTools.buildImage method gaining steam?
So currently at my org we use basically the same scheme I advocated for at a developer conference in 2016 [1]: we just build it all into a gigantic bundled debian package, and deploy that to the host system with apt.
I keep revisiting these various technologies and all of them seem less mature than apt/dpkg, and mostly in service of features and capabilities which don't apply to my particular needs. Obviously I'm a bit of a niche case, and my needs aren't everyone's, but of all of them, nix is the one that seems to be doing the most that is truly interesting and different; being able to send a new nightly build out to users with requiring them to re-downloading an entirely new asset would be a major win.
That aside, there is one slight possible pitfall in discarding one tech because it is less mature than the other: if we assume maturity only increases with time, the oldest product (e.g. apt/dkpg) will always be “best”. Made a note for myself to prefer “not mature enough for my needs” over “less mature than X”.
There's also the ecosystem benefit of having loads of helpers and supplementary tooling applicable to the formats— even stuff like having first-class support in proprietary binary stores like Artifactory, vs Nix where it's basically a shrug and "well... it works with any WebDAV server, so take your pick I guess?"
The main complaint I have overall with Apt is the reliance on postinstall scripts, which means that even if you download and extract your packages in parallelized blocks, you still have a long serialized step when every single package needs to spawn a shell and run arbitrary commands, even if in most cases, the commands actually originate from a semi-declarative format (debhelpers, either invoked explicitly from the rules file, or implicitly by the presence of a corresponding debian/xyz file in the metadata). Anyway, if it were possible to somehow flag packages as atomic or configure-less, it might be possible to significantly speed up these operations, especially in environments like CI where you have everything mirrored in-network or possibly even on-machine so the overall install time is dominated by the package configure step.
Docker, is simple enough to get even the most inexperienced developer going in a very short time.
Docker is the best medium for distributing - A static file is far easier to share / distribute.
Cross-platform - You need an arguably complex and unstable Linux interface to run Docker images, cgroups et al
Sandboxed - security claims about Docker have always been controversial. Simple Unix/BSD constructs like chroot/jails are far simpler and they are reliable
Version pinning - a binary can embed a version and you can stick with it
Reproducible - Everyone gets confused about Docker image checksums. `sha1sum static-binary` is far far simpler.
Minimizes global state - wouldn't be a problem is people built static binaries.
But if none of those are a requirement for your use case (or you have workarounds), I agree that statically linked binaries can be a nicer solution than Docker.
How do you distribute it then? Let's assume your statically linked binary contains both closed-source code and GPL/LGPL code.
> One of the crazy one executable docker containers strikes me as one work to whatever extent a static linked binary is.
I'm not a lawyer, but that's not my understanding.
A docker image is a glorified collection of files with some metadata, just as a tar file is. I think it's broadly agreed that you're allowed to distribute a tar file with unmodified LGPL dynamic libraries in it without having to open-source all of your code, and I think docker images are treated the same way?
You distribute a script and whoever runs the potential violation assembles it themselves, like with zfs on Linux.
I don't really get the demarcations typically made since a proprietary media could conceivably be as hard to pull apart as using linking tools to break apart sections of a static binary again..
I get the general sense that people work around examples of what one interpretation says isn't allowed without getting many opinions on the work around.
https://softwareengineering.stackexchange.com/a/167781
> If the program dynamically links plug-ins, and they make function calls to each other and share data structures, we believe they form a single program, which must be treated as an extension of both the main program and the plug-ins.
To extend this to archives of independent programs, they would be loosely bound, and therefore not form a single program. A docker container that exists only to package up libraries some executable is using would be closer to a single program than a collection of independent components.
Ramping new developers and environments is incredibly trivial with our stack, and we do not rely on any containerization tech. Just .NET Core, visual studio, Git[Hub] and SQLite.
What is unstable about it? As far as I can tell, only the Linux kernel interface is needed, and keeping that stable is an explicit goal of the kernel.
Does glibc work with static linking these days? My understanding was that even with statically linked glibc, things tend to break when the host system has a different libc / a sufficiently newer glibc.
Also, how do you do OpenGL/Vulkan/etc statically? x11docker handles them more-or-less fine, but I'm fairly certain the GPU gods send you to Tartarus if you start trying to statically link in various vendors' libGLs...
> Simple Unix/BSD constructs like chroot/jails are far simpler and they are reliable.
_Fully_ agree about jails, especially nice since they're persistent. Though, for an X11-using application, I think you're screwed any way it comes out, since afaik there's no permissions difference between being able to create a window, and being able to steal keystrokes + send keystrokes to a terminal. Maybe the Qubes people have something?
As far as security on end user machines goes, there's no reason to use Docker over Podman, except for cases where one needs to run docker in docker, which is a farse in and of itself.
I did a project last year at a company that had dockerized their build, CI, and CD infrastructure. They had dozens of git projects with make files that triggered actions using docker. It was great. No need to install anything complicated; just works everywhere with just a minimum of scripts installed from a single internal repository. They did some nice hacks to work around some of the things mentioned in the article. Including using virtual box on macs to work around the filesystem limitations. This really becomes a show stopper for large complicated builds that are very io intensive.
Virtual Machines. VirtualBox in combination with vagrant works reasonably well cross-platform.
vagrant [1] provides this feature based on my understanding.
In my experience that is entirely and ludicrously false. That is merely how a lot of unix software does things by convention, but there are a lot of ways to make even poorly-thought-out unix software behave as a self-contained entity.
In contrast, most linux package managers seem to make a big deal about doing all of this slightly different on just about every linux distribution and even between different versions of the same distributions. The fact package managers exist proves my point: deciding which files go where is a big deal and there seem to be an awful lot of opinionated package managers making different choices here. Whatever standards and conventions exist here seem to leave an awful lot of choice and wiggle room.
Consider that "package managers" only really existed in the UNIX world until relatively recently. Other OSs just didn't make everything so complicated to begin with.
"which files go where" just isn't really a problem if you don't make dependencies some third party's problem.
However, I don’t agree with using it for sandboxing for security. Especially if you are giving it access to the X11 socket, as the author does in the examples. It might not be obvious, but the app will have access to all your other open windows.
Fun fact, most docker hosts will allow access to all your files anyway! (specially true on docker for mac, which all the cool kids(tm) here are using). Even if you restrict container host-FS access to a source repo dir, mind rogue code changing your .git hook scripts in there or you might run code outside of the container when committing ;)
Another slightly relevant fun fact, USB is a bus. That means that any device can listen in on any other device. And USB access is given by default to some X-enabled docker (--tty something), and to most virtualbox machines (including the hidden one running the fake docker linux host on docker-for-mac), and more recently Google-Chrome. ;)
docker-for-mac does not use virtualbox.
Linux and windows:
https://docs.microsoft.com/en-us/dotnet/architecture/moderni...
I use a bash alias for each common application. For example, "jup" launches a jupyter notebook. I have a container with the Python package "Black" which runs using a git hook to clean up my code prior to commits.
On the other hand, you get some benefits from installing black through docker rather than through the system package manager: is is completely isolated from the host and the only way to break black is to update it, changing anything in the host will not break your black install.
I am not sure either what you mean by "just configure your environment properly", but I am going to assume you mean installing black under a virtual env or equivalent? Then it is also annoying for different reasons: you must reinstall it once for each project, updating python to a new (major) version breaks your formatter, you cannot move the env around, to name the ones that come on top of my head.
You can install it either in a venv outside of all projects or even "pip install --user black".
> updating python to a new (major) version breaks your formatter
Uninstalling the old version breaks the formatter. Installing a new one does not. Either way, with asdf, pyenv, and others you can keep all relevant version around.
> you cannot move the env around
Sure you can. Use "--relocatable"
I agree docker may be nicer if you're not working day to day in python... but in that case black is not a great example. You probably want to have a specific version bound to the project so everyone uses it (including the ci platform).
> But in that case black is not a great example. You probably want to have a specific version bound to the project so everyone uses it (including the ci platform).
That's what I do in my open source work, the CI will run the formatter and commit the result back in tree so it forces everyone on one version. At work this is handled by the developer tools' team.
> Uninstalling the old version breaks the formatter. Installing a new one does not. Either way, with asdf, pyenv, and others you can keep all relevant version around.
Yes but at some point you are basically making ad-hoc containers right? So why not use the generic one?
Why would a system package, packaged by experienced maintainers, randomly break?
There were a few configuration steps the last time I configured regular black in IntelliJ, but the documentation wasn't too bad to follow.
Would you be willing to share what you do with Jupyter notebooks, your workflow, how you collaborate with your team, your frustrations?
In my opinion, this is a big enough issue as to throw the whole idea of “Docker as cross-platform platform/target” into major question.
When running more than a few containers on macOS, the performance is so bad it becomes almost unusable.
Once you need a hypervisor most of the benefits are gone. But if you need the hypervisor anyway which might be the case for software that doesn’t have Mac builds then it starts to look attractive.
Changing solutions (e.g Docker for Mac -> VirtualBox or Fusion via e.g docker-machine) makes a world of difference.
The idea is that you mount the filesystem and you got all the tools you need, and more, well installed, and that are lazily pulled from the network.
You win on the space side, but you need a bit more of trust.
You can find more info here: http://packages.redbeardlab.com
On the GitHub repo where you can ask for more packages to be installed:
https://github.com/RedBeardLab/packages.redbeardlab.com/
And in this pair or articles for specific languages
Golang: https://redbeardlab.com/2020/12/21/packages-redbeardlab-com-...
And for JavaScript/node: https://redbeardlab.com/2020/12/23/packages-redbeardlab-com-...
If a distro messes up the trustworthiness of an application, they, the big and important company loses clout.
If the application developer messes up, they also lose clout - people may stop using their software.
Chances are, if you're using a third party for a third party piece of software that isn't officially dockerized by the company that developed it, nor a major distro, there's no real backlash if it doesn't work or if they get hacked, etc: "it was a third party trick, so _of course_ it wasn't trustworthy" would be the statement everyone makes.
Debian messing up, or Cisco or Oracle, etc, is a much bigger deal.
Judging from the Git repo containing my dockerfiles, I've been doing so since ~mid June 2018.
I've since automated:
* checking new versions of Git repos, alpine versions, and short crawlers for tools (i.e. I run "perl latest.pl" and a bunch of stuff happens and eventually some dockerfiles might get updated)
* auto-committing any change made from the above step (i.e. ./autocommit.sh) with a meaningful message based on the directory the dockerfile resides, as well as which environment variable containing the version changed
* I use https://github.com/crazy-max/diun/ running on my dokku server to keep up with base images updates (i.e. I get an email in the morning stating alpine:3.12 has been updated or debian:buster-slim or whatever); when a base image changes I have to manually "dp alpine:3.12" to "docker pull" and "podman pull" it; after that, I "make base-images" and my local base images (each coming with a short line to enable a local apt-cache-ng proxy) to also get updated; then a simple "make" makes all of them (docker build -t .... and podman build -t ...)
* Quite a lot of (mostly small) bash scripts to run those images.
As an example, the Dockerfile I use to build hadolint:
FROM local/mfontani/base:latest AS fetcher
LABEL com.darkpan.github-check github.com/hadolint/hadolint HADOLINT_VERSION
ENV HADOLINT_VERSION v1.19.0
RUN curl -sSL "https://github.com/hadolint/hadolint/releases/download/$HADOLINT_VERSION/hadolint-Linux-x86_64" -o /usr/bin/hadolint
RUN chmod +x /usr/bin/hadolint && \
/usr/bin/hadolint --version
FROM scratch
COPY --from=fetcher /usr/bin/hadolint /usr/bin/hadolint
ENTRYPOINT ["/usr/bin/hadolint"]
... and the shell script I use to run it: #!/bin/bash
DOCKER_FLAGS=()
[[ -t 0 ]] && DOCKER_FLAGS+=(-t)
podman run --rm --init -i "${DOCKER_FLAGS[@]}" \
--network none \
-v "${PWD}:/usr/src:ro" \
--workdir /usr/src \
localhost/mfontani/hadolint "$@"
It's not that speedy doing this, but it's... okay: $ hadolint curl/Dockerfile
Took: 0.837s (837ms)If you are running mainstream Linux(Debian-based/Arch-based, probably other), most of the Docker profits can be achieved with already installed and configured systemd and your distro's package manager.
Sandboxed? systemd.
Simple, uniform interface? Your distro has packages, and most likely services that can and should be sandboxed already run in systemd after installation, you can tune unit-file if you want, and systemd has security checker, that shows you what application in the sandbox can and can not do, without proxying things the Docker way.
Versions pinning? Pin version with your package manager. Want multiple versions? Check out DebianAlternatives system.
Reproducible? Fix your build/install configs, not the environment. If it builds on your machine, but not on the other, or run flawlessly on one, but not the other, it means you have implicit dependency on the environment, or wrong dependency versions constraints, which you likely don't know about. If you don't know your dependencies, you are definitely shooting blind.
Minimizes global state. Repo-based distros minimize global state by providing software that has most dependencies compatible, so you can have your minimal state and update your software too. Meanwhile with docker you need all dependencies and hope that they will match between images, so you can save on layers reusage. And if someone decides to update base image, and others dont... well, too bad, you have to have both versions of the base image.
I don't want to put down anyone for writing these blog posts. The idea is nice from a distance, but the reality doesn't work like that.
It's shoehorning something not suitable for this situation. There's singularity which runs on non-root environments, however it's not for desktop systems, but multi-tenant clusters.
I really get frustrated when people advocate expensive abstractions for minimal gains. We can use our processing power much more efficiently while keeping almost the same properties without the costly abstractions.
Piling everything on top of each other to create impenetrable and immutable abstractions is not the way to achieve this. Docker already makes debugging very hard by being immutable and impenetrable as is.
I wanted to also benchmark bocker[1] (docker written in bash) for a baseline comparison, but it no longer runs and I threw in the towel after ~30 mins of tinkering.
Anyways, if you run something like `curl | jq | grep | less` you could be waiting a fair bit for all those containers to start. Package managers are pretty good these days. I think I can trust it to install `jq` properly.
OK I didn’t know about this. Does each one of the piped commands cause the creation of a new container? And why?
Docker is itself a complex build tool which requires a bunch of install steps. If you are going to ship software to end users there is almost always a better way to bundle and ship than send someone a Docker container.
Docker is not a distribution tool, if you are expecting your end users to install Docker, you've already screwed up.
> Downloading a pre-compiled binary is almost like this, except with worse odds. Maybe there’s a build for your architecture. If it was statically linked, you’re golden. Otherwise, use ldd to reverse engineer the fact that you need to install libjpeg.
On Mac and Windows this is almost never an issue. Even on Linux, it's pretty straight forward to statically link your binary if you aren't sure about the environment it's going to be run on. Statically linked binaries are a bit bloated... but not as bloated as a damned Docker image which contains entire dependency trees.
In no case is "Making it into a Docker Image" a simpler/ better distribution mechanic.
A lot of CI systems natively support Docker because of the variety of tools and images you have available to run you CI steps in.
Docker is an excellent distribution tool because it only requires Docker.
You can send someone some python code and ask them to run it, only to find they're missing a bunch of C libs required to build and run the code which is a pain to help them figure out how to solve.
Docker solves many of these problems but has a drawback of being more 'bloated' than other distribution mechanisms. It doesn't make it a bad one though.
You can send someone a Docker file only to figure out they've never heard of Docker. Let alone don't have it installed and aren't interested in setting up a container system to your 500 line Docker script.
All you've done is move your problem upstream. Unless your consumer is a web developer, you are out of luck.
Trying to debug when a core library or header dependency is missing to build your code requires far more skills and differ between operating systems and versions of operating systems.
The problem isn't moved upstream, it's tackled in a very clever and well packaged manner.
For my home server setup, I have docker containers for:
* PiHole
* NextCloud
* Home Assistant
If I had to install each of those manually, I probably wouldn't have installed them. This is especially true of NextCloud, which almost certainly would have required me to learn how to run nginx on my own, install php or whatever application it uses as the middleware, etc.Instead, I configured my DNS and ran a Docker command and was off to the races.
Server-level Open Source Software's installation process is often so complicated and has so many dependencies if it's not something you can get from your repository's package manager, at least in my experience, that docker is almost always the easiest option.
What beats a single command and maybe reading a config on what ports to forward or how to set up your config. And then you get a docker-compose if you want to be creative, yourself.
Unless you're already very skilled at dev-ops, docker is easier.
I have a huge (personal) wiki page for installing and maintaining NextCloud from before they had a decent Docker image. Now I’m content to let it be a black box I don’t have to think about, so I run it in Docker. It saves a ton of time and hassle.
Docker adds a lot of value when devs ignore the best practice advice of putting everything in a separate container. That’s just a package manager with extra steps IMO. It’s the mini distro style containers like GitLab’s that can save you a masssive amount of time.
If you have an app that formats JSON files, are you shipping it as a Docker container? Or a linter? How about a text editor?
The number of applications where it makes sense to bundle them as VMs is quite small.
Yes, you may end up making a deb and an rpm but honestly, it's not an earth-shattering amount of work, lots of companies do it, and then the tool will tell the user "requires libjpeg".
Jessie also has a blog post about this [1] from back in 2015. If you prefer video format, Jessie also has a talk at DockerCon SF 2015 [2].
[0] https://github.com/jessfraz/dockerfiles
[1] https://blog.jessfraz.com/post/docker-containers-on-the-desk...
One must basically maintain a (more or less) complete userland environment for every application. The idea of shared libraries is taken to absurdity this way. It would be better to build all static.
Then there is the waste of resources. I'm sure with plenty GB RAM, TB of SSD space and GBit of bandwith available nowadays many people don't notice. But for what?
Yeah, you deff notice if you are trying to scale as cheaply as possible…
The advantage is i don't have to fiddle with containers for every little program, no startup delay for every program, and my base ubuntu install stays clean and stable (which runs a VM and other services so stability is important).
It does have a few warts, the main one being you can't have the entire container root volume be a docker volume, so to install new programs persistently inside the devbox i have to rebuild the image (if they don't live completely in the home volume). But logically its an okay tradeoff because the only way to make the whole environment reproducible is to specify things at build time.
For e.g. zoom and hugo don't really change the filesystem or OS settings beyond the folders they output to. I don't see a reason the have them sandboxed personally.
In the end you decide at what point are you willing to delegate responsibility for things working as they say they should.
Still use Docker alongside WSL2 for purely-Linux stuff like some node.js scripts or python things that don't need GPU.
Just look at something like this which tries to compile C++ for VcPkg: https://hub.docker.com/r/hripko/vcpkg/tags?page=1&ordering=l... (the Linux images are all < 300Mb, but the Windows one is 5Gb compressed).
Or use named volumes. I'm running a dockerized WordPress dev environment on my MacBook with average TTFB's of 40 ms.
There was a very in-depth thread on the docker forums where the devs explained why there was such a huge performance penalty. IIRC it was due to all the extra bookkeeping that had to be done to ensure strong consistency and correct propagation of file system events between the virtualized docker for mac environment and the host file system.
The test suite would run integration tests that performed a lot of npm/yarn operations which meant lots of disk IO.
And how do you access these from your host with high performant IO?
Many developers would consider this Ubuntu 20.04 install that I run on a 9 year old computer to be an accident waiting to happen. But I'm finding this Ubuntu release to be extremely stable even at 3000+ packages.
I've been using debian for a long time and would prefer not to have to switch OSs but the idea of using nix for having full control over my package graph is very tempting.
I've been somewhat procrastinating on trying nix as I've heard GNU Guix has a similar feature set and haven't been able to decide on which one to dive into...
My ideal setup would be to just be able to run a single shell script that configures a new machine to the exact state of all my other dev machines. I have a shell script that somewhat does this but it's not completely unattended and still requires a lot of manual config for certain steps.
Aside from that my main use case is being able to easily share a dev/build environment with others for ensuring that they can compile a certain project exactly as I do. For now I just use docker but it's frustrating not having explicit control over the layer cache and being able to tell it what to cache and what not to.
My experience: no. I've been very happily using Nix on debian for years. Debian gives me my boring desktop apps (I have very very boring tastes), almost all my tinkering and development starts with drawing all the prerequisites from Nix.
The only times you get weirdness from using non-NixOS linux are things like running opengl apps, which simply require a wrapper script like `nixGL` to function properly.
And darwin Nix is just about the only thing that makes macos a bearable platform for me.
Perhaps Nix might help here, but I've never used it, so can't say for sure.
Worth pointing out that there is an incubating CNCF project that tries to solve this problem by forgoing Dockerfiles entirely: Cloud Native Buildpacks (https://buildpacks.io)
CNB defines safe seams between OCI image layers so that can be replaced out of order, directly on any Docker registry (only JSON requests), and en-mass. This means you can, e.g., instantly update all of your OS packages for your 1000+ containers without running any builds, as long as you use an LTS distribution with strong ABI promises (e.g., Ubuntu 20.04). Most major cloud vendors have quietly adopted it, especially for function builds: https://github.com/buildpacks/community/blob/main/ADOPTERS.m...
You might recognize "buildpacks" from Heroku, and in fact the project was started several years ago in the CNCF by the folks who maintained the Heroku and Cloud Foundry buildpacks in the pre-Dockerfile era.
[Disclaimer: I'm one of the founders of the project, on the VMware (formerly Cloud Foundry) side.]
In particular the out of order layer replacement. I'm interested in switching to Buildpack for the images I maintain for my home cluster. Would make upgrading my base image so much simpler compared to rebuilding all the other images! I read a bunch of docs/articles since reading your comment yesterday but couldn't find any mention of this, or better yet an example. Are there some docs I missed? (I didn't look into the spec.)
A few tips on rebase:
(1) If you want to rebase without pulling the images first (so there's no appreciable data transfer in either direction), you currently have to pass `--publish`.
(2) If you need to rebase against your own copy of the runtime base image (e.g., because you relocated the upstream copy to your own registry), you can pass `--run-image <ref>`.
I also don't really care for the "we have so much space now" argument. I certainly don't, I don't put in expensive 2TB SSDs in my laptop because I don't need them, and I don't want half of my 500GiB disk to be taken up by giant blobs of unoptimized docker images for the same reason that I don't want to run 10 copies of chrome at the same time to use a text editor, an email client, a web browser, a media player, a debugger, the thing I'm writing, three chat clients, and a partridge in a pear tree. I have extra space and extra processing power on my computer so that a: I can use it for the things I actually want to use it for and b: so that I have a snappy machine which can take an unexpected load (be it disk load or processing load) without problems.
However, Docker is like a kitchen machine: good for getting something done in a particular way (stand mixer for kneading dough) but less useful for understanding what’s actually happening (when is pizza dough ready?)
On Linux, picking apart LXC stuff at the command line is very much worth your time if you’re interested in how things work. (Not LXD though which is more useful as a tool than as a teaching aid.)
If you have some IPv6 allocation to play with you can make some interesting infrastructure and have it do something useful on the Internet without relying on the lxc-net crutches of an automatically built bridge with NAT. It all feels very well designed as a bag of tools to let one make things rather than a complete system that guides you in only one particular direction.
I use stand mixers and Docker all the time; sometimes it’s fun to get into the details too.
https://github.com/casidiablo/macondo
I don't even know what it is I built, but it has been useful in some contexts.
It basically allows you to easily wrap and distribute scripts (or more complex apps) that have specific dependencies that might not always be installed in the host. It does so by wrapping the script in a docker image.
It also automates the annoying part of docker: mounting local paths for apps that need to interact with the host's file system.
I need to write a blog post on this if anything to gather feedback. I'm still not 100% sold on the idea and there are some edge cases. Still, a fun experiment.
So ultimately the recommendation is to install Linux and use it as my daily driver. But Linux does not run on my machine yet.
What's really needed are better/more articles about how to clean off the cruft left by docker after running your builds locally. Especially when you do things like update your base image versions from node:12 to node:14 etc.
A large amount of inter-connected projects for a SaaS-like service, all mounted inside the docker container, with some very specific libraries pre-compiled in the docker image, that used to take 30+ minutes to build.
And now they're just in a docker image, that will work on linux and mac and I guess windows too.
Nice idea in theory.
No, thanks - I shudder at the existential nightmare at being forced to do all development via Docker. Use Docker for deployment, exit the VM and get back to your work.
Docker has tradeoffs.
Would be good to hear her thoughts a few years on
As workaround for broken dependency management, not really.
Native binaries can easily do with static linking (fix glibc or replace it with proper libc like musl).
Other platforms are doing quite fine without containers.
But in the Linux world no one has the intention to fix the package distribution problem except for that AppImage guy.
It's not really portable, it only runs decently on Linux, on other platforms it's a kludge.
On a mac this is a complete deal breaker for anything you need to run frequently. It is more than slight.