Having recently delved into containers, I don't like this quote. It's not wrong, but I think it's also misleading. In particular:
> Whatever is the underlying execution system
Docker can only execute containers on Linux. Indeed, Docker itself is essentially a UI and management tool for Linux kernel features (chroot, bind mounts, etc.). Whilst Docker provides software for other platforms, like macOS, those are just UIs bundled with a VM/hypervisor running Linux; and it's that Linux system which runs the containers.
This may seem pedantic, but I think it's important to understand (a) what's actually happening when we run a container, and (b) what we can/can't do with such tools. For example, I was confused why my containers couldn't run macOS binaries: the whole point of containers is that they run inside the host OS, unlike VMs which have their own OS. That's certainly true, but I didn't realise that when running Docker on macOS, the host is still a VM running Linux!
> Docker can run it
Again, pedantic but important: Docker is a UI and management tool for containers, which are ultimately run by an underlying Linux system (usually via `runc`, or something compatible). I find this important since "Docker" provides much more than just running containers, e.g. it has Dockerfiles, images, registries, etc.
If someone just wants to run a container, they don't need Docker; and I've found Docker to be a more complicated and convoluted way of running containers than, say, OCI tools.
For example, I recently wrote an AWS Lambda function which needs to run from a custom container (it bundles some tools which are larger than Lambda's 50MB code limit; whilst containers have a 10GB limit). I originally did this with Docker, which required:
- Building an 'image' containing the software (I actually did this with Nix, rather than `docker build`, for reasons I'll give below)
- "Loading" that image into Docker
- "Tagging" the image with the URL of an ECR 'image repository' (we could include this tag during building, but I find this way keeps more distinction between 'building' and 'deploying')
- Running `aws ecr get-login | docker login` to "log in" to Docker. This seems ludicrous to me; I hear it's something to do with dockerhub compatibility or somesuch; which is still silly, since we're not using that.
- Pushing to ECR using Docker
In contrast, I've now switched that project to use OCI images, which only requires the following:
- Building an 'image' containing the software. This is just a .tar.gz file, plus some .json files which specify the "EntryPoint", the SHA256 of the .tar.gz, etc. (all easily created with bash + tar + jq)
- Uploading the image to ECR. This can be done with `aws ecr` shell commands, but they're quite low-level (e.g. files are uploaded 20MB at a time, which needs a loop) so I did this in Python using boto3. Note that the 'tag' is still needed, but it's just an argument to the 'ecr.put_image' function.
> the exact same code, byte by byte
This is my main problem with the quote. Whilst it might theoretically be possible to use Docker "properly", it seems to actively encourage awful practices.
For example, to run "the exact same code, byte for byte" we would need to actually get those bytes in the first place. Whilst the underlying container is simply a directory (known as a "bundle", which `runc` actually executes), Docker abstracts over such bundles: first we tar them up into "layers", which we then list in a JSON file called an "image", which we then "push" to an "image repository". To run a container, we "pull" its image from a repository, which we reference by a "tag" (essentially a filename).
The de facto tag is "latest", which appears in all sorts of documentation and tutorials. Tags may get replaced with new images at any time by another push, which makes the idea of "the exact same code, byte for byte" rather misleading.
(Thankfully some repositories, like ECR, allow 'immutable tags', which can never be overwritten. Rather than 'latest', I use hashes of an image's content as its tag; so there's never a conflict when I push a new image.)
Next, we can ask how those bytes come into existence in the first place. In the world of Docker, we don't put files into a .tar.gz; instead, we download an entire Linux distribution, then we run a bunch of non-deterministic package management commands like `apt-get update && apt-get install -y foo`, resulting in a massive, unreproducible binary blob.
Whilst it can be argued that Docker will run 'the exact same code, byte by byte', it can also be argued that coin tosses are deterministic due to Newton's laws of motion (e.g. a reliable Docker system is illustrated in Figure 1 of http://statweb.stanford.edu/~susan/papers/headswithJ.pdf )