What Exactly is Docker?
medium.com
medium.com
> RUN apk update
> RUN apk add chromium chromium-chromedriver
These kind of poor examples lead to a huge amount of waste when using Docker because people learning it are not taught about how layers interact, leading to ridiculous things like 'COPY' followed by 'RUN chown'.
Layers are a _core_ part of how Docker works, why not make this example "correct" by doing:
> RUN apk --no-cache add chromium chromium-chromedriver
Then just a comment like this: "it's important to group layers if possible to reduce your image size. By using `--no-cache` apk will update at the same time as installing".
I think this post is attempting to oversimplify, and is doing a bad job of teaching in the process.
Try doing anything hardware-network-topology relevant stuff without `network=host`, like running a ROS cluster with direct ethernet connections between hosts and cameras.
It's definitely "on top" in a way bare-metal isn't. Same way that a vm is "on top of" kvm even though the cpu supports hardware virtualization.
Nvidia-docker is another abstraction. It's a pass-through. It's not the same as bare-metal. Certainly feels the same 98% of the time, but in that 2% it is very much apparent there is a layer in between.
> RUN echo something > some_file
> RUN rm some_file
Neglecting to even mention it leads to issues down the line because people don't get the fundamentals. There are a lot more confusing things than this when working with Docker.
Really nobody really cares about container sizes except for their own efficiency. We run thousands of containers. Many of them tens of GB. They could all be optimized to a small size but it really doesn’t matter and is not worth the engineering time most of the time.
And yes, if you're adding 5+ GB to each image build through something silly like a 'RUN chown' it's worth fixing. Maybe not complex refactoring, sure, but we are talking about the basics here.
This Docker and Kubernetes hype and the fact that few people understand and care to implement these technologies correctly fills Amazon's and Google's pockets with money that could be spent elsewhere.
Learning layers is not critical to get up and running with Docker. You can learn it easily afterwards, but on the very first introduction it's just unnecessary cruft that distracts you from your goal.
More importantly, you can actually use Docker in daily practical applications your whole life not caring about layers. Think about it: why should you care about layers if you don't care about build size or build time?
At this stage, caring for layers is just premature optimization.
RUN cmd1 && cmd2 && cmd3 && cmd4
maybe with brackets:
# 1 layer
{
RUN cmd1
COPY someshit
RUN cmd2 -lots -of -long flags
...
}I've not used it but it's on my list to try out, the explicit caching I think ties in with what you're after:
* Requires no elevated privileges or containerd/Docker daemon, making the build process portable.
* Uses a distributed layer cache to improve performance across a build cluster.
* Provides control over generated layers with a new optional keyword #!COMMIT, reducing the number of layers in images.
* Is Docker compatible. Note, the Dockerfile parser in Makisu is opinionated in some scenarios. More details can be found here.
Alternatively you can use buildah [1] to build your images via a script, but you'll lose the cache docker puts in place.
Thanks a lot for the feedback. I believe simplicity at this phase of learning is much more valuable than efficiency, that's why I haven't tried to produce the most efficient images as examples. Since your suggestion also simplifies the Dockerfile itself, I have replaced the previous one with your suggestion, thanks!
I think a lot of it has to do with the experts accidentally talking past beginners, missing a lot of the basics before getting into teaching abstractions.
It also reminds me of the feeling I'm experiencing now about learning Elasticsearch. I'm amazed just how few JSON examples I can find online for the API. It was amazing how much it helped for a peer to say, "an index is a table, a document is a record and it's kind of like monogdb."
Furthermore this all reminds me of wrong atomic models in high school. Please just teach me a really simple but wrong explanation then slowly work out the details.
If you're a programmer, you most probably have heard about OSes and VMs before, one way or another. Then, the analogy is a bit easier for the beginner to understand. Hence the "slightly wrong but somehow still manages to be a bit informative"-summary Waterluvian gave can teach the beginner the first steps.
Maybe it's just me. I love teaching and learning by starting with a common but wrong model and learning piece by piece what's different until its an evolved, more accurate model.
Not really, no. Software in a container directly talks to the kernel of the host using the normal APIs the kernel provides. A container does not contain a kernel!
Running an Ubuntu user land on top of some generic other Linux kernel will for almost all practical purposes feel the same as running real Ubuntu.
That kernel is what is shared between different containers. You can have 5 copies of Ubuntu (or Debian, or Mint) in separate containers, but they're all using the same kernel. This is what makes containers much more efficient than multiple full VMs.
They are just applications running on top of the very same kernel you would run any application from.
The kernel just provides enough layers of separation on the necessary structures.
There's plenty of material on it; if interested I'd suggest to start to study by kernel namespaces as @mav3rik pointed out.
Truly understanding the Linux kernel namespacing features that Docker is built on certainly takes more effort to learn, but I disagree that it's irrelevant to anyone not working at Docker. Understanding your tools makes you a better developer, and allows you to understand what is and isn't possible.
I think it is a fair summary, because the implications of running in Docker are often the same as running a real VM.
To exemplify, a common mistake is to run a server as the entrypoint of the docker container. The server will then run as PID 1, with all the implications and responsibilities an init process has. This causes a lot of subtile and annoying problems.
Unless you are writing kernel drivers or you need to tweak certain kernel parameters, running in a Docker container is very much as running in a VM. With the same responsibilities, e.g., setting up an init process.
This is exactly what Docker containers AREN'T.
Don't think of containers like tiny VMs. Think of them as processes with additional isolation from each other.
Under the hood they are simply a process isolated with namespaces, but their behavior on the outside feels like getting a VM.
Unless one is 31337-rockstar-ninja-IQ150-programmer, the uptake of new skills is painful and takes time. The pedagogical process is important. It's totally OK to not know how anything works as long as you acknowledge that and are willing to learn.
I'm mostly an ops person in outlook, although I was a programmer for over a decade and write my share of code. The software-ization of infra is a good thing; unlike some ops folks, I don't find it threatening, I find it cool.
What I find dismaying is some of the horrific stuff I catch before it gets to production, and what I do find threatening in a different way is worrying about what I miss.
Moving OS-level dependency management in to dockerfiles? Great! How are you scanning that against CVEs? If you don't know you're not running a different kernel, I have a sneaking suspicion you're... not. And so on down the list of traditional ops practices I frequently discover developers had no idea even happened.
There is nowhere in the docker-noob's workflow where the huge distinction between a docker container instance and a super light VM actually matters.
Source: I knew nearly nothing about docker a year ago. And now I'm one of my company's go-to people for docker questions. The pain is still fresh in my mind and I see noob-pain every day :)
Aaaand that off-the-cuff statement clears up the difference between them that I've been missing for months. Containers really do just look like a VM, if you don't know the implementation details and are just using them.
You probably didn’t mean that they literally emulate an instruction set, because most modern hypervisors don’t emulate most instructions.
The idea of a VM is so messy these days, it's best to define what kind of VM are you talking about first. (language? system? foreign hardware?)
That's true, but as VMs also isolate processes from each other, this can confuse others. Let's be more exact:
A VM = isolated kernel and userspace.
- Advantages: security
- Disadvantages: speed
A container = isolated userspace only. Shared kernel.
- Advantages: speed, size
- Disadvantages: security
Both can have pre-created images to avoid the problem of recreating workspace. Docker did not invent orchestration.
Nuh? Of starting, sure. Of running? A margin of error.
- Containers in this case: userspace on kernel on hardware
- VMs in this case: userspace on kernel on virtual hardware on kernel on hardware.
There are very fast VMs (the Firecracker MIcroVM) but they're not in popular use.
VMs are qemu-system-x86_64 -machine accel=kvm or whatever is the platform specific way to run a VM under a hypervisor. This has been the case since hypervisors showed up on the scene, aka right after what was the predecessor of Virtuozzo
The overhead of running a VM absolutely changes depending on whether it is a type 1 hypervisor or a type 2 hypervisor or if the runner/monitor supports acceleration.
You can run a VM with QEMU without accel=kvm and it will be dog slow. Try it.
For CPU/memory, yes.
For network, possibly (lots of iptables re-writes)
for disk, not so much, docker(depending on what driver you have) suffers from write amplification.
If I gave a pre-Docker era engineer two terminals, one host, one in a container, they'd probably quickly figure out they are on the same physical machine. But unless they were super savvy at some very specific stuff, they'd probably conclude it was just a VM.
Every dev I've introduced to docker workflows basically treat it like a vm. The only ones that didn't already had experience with cgroups/chroot and related tools.
It's honestly better to think of Docker as "VM, but actually you use the same kernel" than "processes that are isolated".
Other poster compared it to electron models. Yes, "balls in orbits" is wrong, but kids are already familiar with planetary orbits, and that mental model works until basically Organic Chem II.
1. For software, open source or not, that is driven by a commercial company, it is often not in their best interest to tell people what it actually is, that they are selling. They care more about conversion rates than educating. E.g. notice how the docker.com website doesn't seem to explain at all what these "Docker containers" are, instead throw around phrases the marketing department came up with: "Docker is the de facto developer standard for building and sharing apps that enable simplicity, agility and choice [...]".
Doesn't mean the tech is good or bad. It can just make it obnoxious to sift trough the propaganda to get the info you want/need.
2. There are plenty of developers that can't be bothered to obtain a deeper understanding of the things they are working with, beyond the getting started tutorial. Which is fine in the beginning, but eventually one should get past that.
You're right. It's only slightly better than zombo.com (https://html5zombo.com/)
Would you prefer to know what docker containers are, or what they provide? For example would you rather the first explanation started "Containers allow you to run applications in an environment with its own filesystem and network, isolated from the resources of the main system", or "Containers are a standardised abstraction layer over linux namespaces and resource use restrictions, with common interfaces for network and volume management"? And would you stop using containers if they provided the same benefits, but technically in a different way?
Sad to see he gave up before he started, and instead of explaining what Docker is, went off into the docker and docker-compose CLI commands. What an opaque explanation too :(
Docker is hard to explain, and the official documentation won't help you understand it. I'm sad that so few people are self aware enough to combine just the right, minimal depth of concepts about kernels, operating systems, systemd, namespaces, and the fact this all only works on Linux, to make a truly approachable explanation. Most developers are really bad at teaching, they only describe things they already know, vs actually trying to teach something.
Docker runs on Windows. Windows containers run on Windows. Linux containers run on Linux as well as Windows. (To run Linux containers on Windows a small Linux kernel is run inside HyperV.)
You can think of Docker containers as VMs, except instead of running its own copy of the OS it runs directly on top of the host machine's OS.
I understand some bits and pieces about how containers work internally on different platforms, I also have a draft article that goes into depth with these, but the goal of this specific article was to give some general information as a basis for people to get started. You don't have to how namespaces work on Linux in order to build a Dockerfile for your project and share it with your colleagues. You don't have to have the most optimal image in order to replace your `git pull` on your server with a simple Docker Compose setup. My goal was to give high-level overview of these technologies and include actionable examples so that the improvements I believe these technologies bring can be applied immediately by the readers.
> Most developers are really bad at teaching As per this one, I am honestly sad to read this, though I am glad that you gave this feedback. I had a really hard time a couple of years ago when I was starting to learn these technologies, and I remember very vividly that everything I read was very abstract and lost in detail, which made me feel very dumb. I personally believe that any complex topic can be explained to a layman if the teacher knows the topic well enough, and I wanted to attempt to build a document that I wish existed at the time I was trying to learn these stuff. I was thinking I did a decent job with this article, but clearly there are some short-comings in terms of the content and the way I explained things. If you have any suggestions, please feel free to give suggestions so that I can make improvements to the article and keep it as a living document, I am genuinely interested in improving both my skills and this article itself.
Again, thanks a lot for the feedback.
Please watch this awesome presentation: https://www.youtube.com/watch?v=zGw_xKF47T0
On Windows it is also possible to run Windows containers (no Linux at all) that use the Windows kernel and run Windows programs.
If you have suggestions to improve my point above in the article itself I'd be glad to take that input and incorporate it into the article itself, feel free to write here or reach out to me via email in my bio. Again, thank you for the feedback.
you do not. you may also install Podman. Docker does not "own" containers, there is an open standard for containers that any vendor may implement.
Also I'm disappointed that the good old chroot is not mentioned, or the BSD jail system.
I think this is one of the worst "Best Practices" ideas that are parroted by people who haven't thought deeply about the issue. It's really a bad legacy from the era when most software was actually distributed. Now that most software runs in environments that are controlled by the same organization that developed the software, the principle is far less valuable.
Nowadays, most software should have most of its configuration information - paths, DB URLs, HTTP endpoints, etc - hard-coded into it. This strategy follows the "convention over configuration" philosophy, and it gives you a range of benefits. First of all, you can run tests on your config to make sure everything is working properly (check various files are present, do a SELECT * LIMIT 1 from DB tables, etc). You can catch config errors at compile time, eg by using enums like prod/dev/qa to represent environment names. And it prompts you to apply a refactoring mindset to your config - when you notice that your config code is repeating itself extensively, you'll realize this and be able to take steps to refactor, standardize, and simplify the config.
Do you have references to companies that are developing software this way at scale?
My current solution is to give each user their own docker daemon running in a dedicated virtual machine... Do you have a better solution?
For 99% of the world, the answer is "a file format like .tar.gz except composable".
These guys are really missing their target audience needs by a mile.
The goal of Nix and Nixpkgs is to have effient recipes for building everything ever, in all configurations. The docker ecosystem could never get there.
Now containers do make sense for deployment, but that has little to do with docker, as those docker replacements for kubernetes demonstrate.
Packaging is just one component of a container, and it actually works quite well in docker, since it feels like you’re packaging up the entire OS.
Instead of a cli tool.
Because there are already standalone docker implementations that implemented with completely different technology. Just like docker on windows (the one runs exe).