LXC vs. Docker
earthly.dev
earthly.dev
Nothing stops one from running Docker inside LXC. For development I usually just make a dedicated priviledged LXC container with nesting enabled to avoid some known issues and painful config. LXC containers could be on a private network and a reverse proxy on the host could map to the only required ports, without thinking what ports Docker or oneself could have accidentally made public.
But even if they had, the RAM snapshot needs to be written, but without freezing the container. I would appreciate an option when I could ignore everything that was not fsyned, e.g. Postgres use case. In that case the normal ZFS snapshot should be enough.
The state of a container is not part of the snapshot (just checked), as it is really hard to capture the state of the container (CPU, network, all kinds of file and socket handles) and restore it because all an LXC container is, is local processes in their separate cgroups. This is also the reason why a live migration is not really possible right now, as all that would need to be cut out from the current host machine and restore in the target machine.
This is much easier for VMs as Qemu offers a nice abstraction layer.
We've pretty much stopped using LXC containers in Proxmox because of all the little issues.
Things that got me recently, just off the top of my head:
- htop sees memory usage incorrectly
- Docker tries to use overlay2 on ZFS which fails (I think, I needed to create and mount an ext4 volume for reasons)
- Hashicorp Vault needed disable_mlock because I believe LXC blocks the syscall.
On the other hand, I like my Samba file server in a container though, because it is much easier to share storage from the host into LXC than a VM.
LXC and VMs both have pros and cons.
Reminds me of (now defunct?) flockport.com
They had some interesting demos up on YouTube, showcasing what looked like a sandstorm.io esque setup.
Proxmox, a few LXC, each with their own containerisation running.
LXC can be directly compared with a small, and quite insignificant, part of Docker: container runtime. Docker became popular not because it can run containers, many tools before Docker could do that (LXC included).
Docker became popular because it allows one to build, publish and then consume containers.
At least they mentioned it in their apple:oranges comparison.
there's also the global namespace thing. "FROM ubuntu:18.04" is pretty powerful.
I run a proxmox server with LXC but if I could use a dockerfile or equivalent, my containers would be much much more organized. I wouldn't like to pull images from the internet however.
Reproducible? No. Most Dockerfiles are incredibly unreproducible.
Yes, baking doesn't work the same at sea level vs say denver colorado @ 1 mile, but pretty much you can share your recipe with friends and save them lots of time.
It is possible to get to reproducibility though, by involving lots of checksums. For example you can use the checksum of a base image in your initial FROM line. You can download libraries as source code and install them. You can use package managers with lock files to get to reproducible environments of the language you are using. You can check checksums of other downloaded files. It takes work, but sometimes it is worth it.
Docker was good enough, simple enough, and got the job done. It won for good reason, but the next generation should learn from it's mistakes.
True, but Docker is an awful choice for those things (builds are performed "inside out" and aren't reproducible, publishing produces unauditable binary-blobs, consumption bypasses cryptographic security by fetching "latest" tags, etc.)
Docker supports multi-stage builds. They are quite powerful and allow you go beyond the "inside out" model (which still works fine for many use cases).
> ...and aren't reproducible
You can have reproducible builds with Docker. But Docker does not require your build to be reproducible. This allowed it to be widely adopted, because it meets users where they are. You can switch your imperfect build to Docker now, and gradually improve it over time.
This is a pragmatic approach which in the long run improves the state of the art more than a purist approach.
I though Dockerfile ensured that builds are indeed reproducible?
Lots of docker build scripts have the equivalent of
date > file.txt
or curl https://www.random.org/integers/?num=1&min=1&max=1000&col=1&base=10&format=plain&rnd=new > file.txt
Buried deep somewhere in the code.But yeah I don't see any reason why you couldn't theoretically make a reproducible build with Docker.
Buildah from Red Hat has an argument to set this programmatically instead of using the current date, but AFAIK there's no way to do that with plain old docker build.
If you wanted to guarantee reproducibility, a hypothetical purely-functional build system could do that. I personally don't think that's necessary though, or that it would be worth the requisite trade-offs.
> I [thought] Dockerfile ensured that builds are indeed reproducible?
The example straightforwardly disproves that.
That said, the above example isn’t specific to Docker.
This is probably a good argument because of how hard it is to do anything in a reproducible manner, if you care even about timestamps and such matching up.
Yet, i'd like to disagree that it's because of inherent flaws with Docker, merely how most people choose to build their software. Nobody wants to use their own Nexus instance as a storage for a small set of audited dependencies, configure it to be the only source for all of them, build their own base images, seek out alternatives for all of the web integrated build plugins etc.
Most people just want to feed the machine a single Dockerfile (or the technology specific equivalent, e.g. pom.xml) and get something that works out and thus the concerns around reproducibility get neglected. Just look at how much effort the folks over at Debian have put into reproducibility: https://wiki.debian.org/ReproducibleBuilds
That said, a decent middle ground is to use a package cache for your app dependencies, specific pinned versions of base images (or build your own ones on top of the common ones, e.g. a customized Alpine base image) and multi stage builds, you are probably 80% of the way there, since if need be, you could just dump the image's file system and diff it against a known copy.
Nexus (some prefer Artifactory, some other solutions): https://www.sonatype.com/products/repository-oss
Multi stage builds: https://docs.docker.com/develop/develop-images/multistage-bu...
The rest 20% might take a decade until reproducibility is as user friendly as Docker currently is, just look at how slowly Nix is adopted.
> publishing produces unauditable binary-blobs
It's just a file system that consists of a bunch of layers, isn't it? What prevents you from doing:
docker run --name dump-test alpine:some-very-specific-version sh -c exit
docker export -o alpine.tar dump-test
docker rm dump-test
You get an archive that's the full file system of the container. Of course, you still need to check everything that's actually inside of it and where it came from (at least the image persists the information about how it was built normally), but to me it definitely seems doable> consumption bypasses cryptographic security by fetching "latest" tags
I'm not sure how security is bypassed if the user chooses to use whatever is the latest released version. That just seems like a bad default on Docker's part and a careless action on the user's part.
Actually, there's no reason why you should limit yourself to just using tags, since something like "my-image:2022-02-18" might be accidentally overwritten unless your repo specifically prevents this from being allowed. If you want, you can actually run images by their hashes, for example, Harbor makes this easy to do by letting you copy those values from their UI, though you can also do so manually.
For example, let's say that we have two Dockerfiles:
# testA.Dockerfile
FROM alpine:some-very-specific-version
RUN mkdir /test && echo "A" > /test/file
CMD cat/test/file
# testB.Dockerfile
FROM alpine:some-very-specific-version
RUN mkdir /test && echo "B" > /test/file
CMD cat/test/file
If we use just tags to refer to the images, we can eventually have them be overriden, which can be problematic: # Example of using version tags, possibly problematic
docker build -t test-a -f testA.Dockerfile .
docker run --rm test-a
docker build -t test-a -f testB.Dockerfile .
docker run --rm test-a
In the second case we get the "B" output even though the tag is "test-a" because of a typo, user error or something else. Yet, we can also use hashes: # Example of using hashes, more dependable
docker build -t test-a -f testA.Dockerfile .
docker image inspect test-a | grep "Id"
docker run --rm "sha256:93ee9f8e3b373940e04411a370a909b586e2ef882eef937ca4d9e44083cece7c"
docker build -t test-a -f testB.Dockerfile .
docker image inspect test-a | grep "Id"
docker run --rm "sha256:8dd9ba5f1544c327b55cbb75f314cea629cfb6bbfd563fe41f40e742e51348e2"
docker build -t test-a -f testA.Dockerfile .
docker image inspect test-a | grep "Id"
docker run --rm "sha256:93ee9f8e3b373940e04411a370a909b586e2ef882eef937ca4d9e44083cece7c"
Here we see that if your underlying build is reproducible, then the resulting image hash for the same container will be stable. Furthermore, someone overwriting the test-a tag didn't break you being able to run the first correctly built image because the tags are just convenience, so you'll be able to run the previous one.Of course, that loops back to the reproducible build discussion if you care about hashes matching up, rather than just tags not being overwritten.
IMHO, the real 'trick' with Docker isn't really container runtimes, images, layers, etc. It's the willingness to avoid system dependencies in favour of doing everything inside a "container image" (AKA .tar.gz). That would seem crazy to a Makefile writer in the 80s, but once we become willing to do this, the actual technique can be implemented using something like Make (with appropriate use of `./configure --prefix` arguments, 'export PATH=...' commands, etc.).
Sure it would be leaky, inefficient, etc. but as you say, the majority of devs wouldn't mind (just like with Docker).
To be clear, I'm not advocating anyone actually try doing this with Make, or whatever. I'm just pointing out that many of the claimed advantages of containers (in general) and Docker (in particular), like being sort-of isolated, or sort-of cross-platform, etc. are actually completely orthogonal to the underlying technology. Instead, those advantages come from the way they tend to be used.
Unfortunately, some of the downsides also come from the way they tend to be used (e.g. putting an entire OS inside a container, rather than just the intended binary + its deps; or using 'latest' tags instead of hashes)
> It's just a file system that consists of a bunch of layers, isn't it?
Exactly. Hence it's hard to check whether, for example, the bin/foo executable contains a patch for CVE-1234, or whatever.
Compare this to e.g. Maven .poms, Nix .drvs, etc. which tell us what went into any particular artifact.
> Actually, there's no reason why you should limit yourself to just using tags, since something like "my-image:2022-02-18" might be accidentally overwritten unless your repo specifically prevents this from being allowed. If you want, you can actually run images by their hashes
Indeed, this is actually a really nice thing about Docker (which has since been incorporated into OCI). However, the tragedy is that it tends to get bypassed in favour of tags, and more specifically just 'latest'.
For a fantastic way to work with LXC containers I recommend the free and open Debian based hypervisor distribution Proxmox [1].
[0], https://thehftguy.com/2016/11/01/docker-in-production-an-his...
Going forward we built and installed lxd from source.
If enough people were to ever decide to get together and properly fork snapd and maintain the patched version I'd totally dedicate time to helping out.
https://gist.github.com/alyandon/97813f577fe906497495439c37d...
Very annoying.
I just avoid snap and Ubuntu wherever possible now.
Is that the gist of flatpak?
Who even uses snap in production? If I squint my eyes I can see the use for desktops, but why insist on it for server technologies as well?
That all being said, LXD is great way to run non-ephemeral containers that behave more like a VM. Also checkout multipass, by Canonical that also makes spinning up Ubuntu VMs as easy as Docker.
I belive it is something as easy as
lxc launch -vm ubuntu/20.04 myvmAfter that I setup minikube with a Docker backend. It all worked instantly, perfectly aligned with my mental model, zero bugs, zero hassle. Canonical builds a great OS, but their Snap/VM org is... not competitive.
The only gripe I have with Alpine is its installation experience. Like Arch, Alpine has a DIY type installation (but a completely different style). But unlike Arch, it isn't easy to properly install Alpine without a lot of trial and error. Alpine documentation felt like it neglects a lot of important edge cases that trip you up during installation. Arch wiki is excellent on that aspect - they are likely to cover every misstep or unexpected problem you may encounter during installation.
You can control snap updates to match your maintenance windows, or just defer them. Documentation here: https://snapcraft.io/docs/keeping-snaps-up-to-date#heading--...
What you cannot do without patching is defer an update for more than 90 days. [Edit: well, you sort of can, by bypassing the store and "sideloading" instead: https://forum.snapcraft.io/t/disabling-automatic-refresh-for...]
It is much easier to give sysadmins back the power they've always had to perform updates when it is reasonable to do so on their own schedule without having to inform snapd of what that schedule is.
I really don't appreciate being told by Canonical that I have to jump through additional hoops and spend additional time satisfying snapd to maintain the systems under my control.
I think the reason for auto-updates like this is because selling the control back to us is the business plan. It's the same thing Microsoft does.
First they had just moved it to Snap which was not a great install experience compared to good old apt-get, and then all my containers had no IPv4 because of systemd for a reason I can't remember.
After two or three tries I just gave up, installed CapRover (still in use today) and have not tried again since.
It is not good even on single dev machines.
I could not bear the snaps on ubuntu always coming back and hard to disable on every update, I gave up and just switched to arch and happy to have control on my system again.
I had a lot of crash running on Ubuntu when running huge rust based test suite doing a lot of IO (on btrfs), never had that issue on arch. not sure why, not sure how I can even debug it (full freeze, nothing in systemd logs) so I guess I just gave up.....
LXD is my perfect fit in this scenario: trivial to install on top of Nixos, and once running, allows for launching some minimal development instances of whatever distro flavor of the day in a few seconds. Persistent like a small VM, but booting up within seconds, much more efficient on resources (memory in particular), and - unlike docker - with the full power of systemd and all. Add tailscale and sshd to the mix, for easy, secure and direct remote access to the virtualized system.
However, an exciting thing to me is the Cambrian explosion of alternatives to docker: podman, nerdctl, even lima for creating a linux vm and using containerd on macos looks interesting.
That said, you're going to be doing some plumbing for things like wiring your services together (Fabio/Consul Connect are good choices), detecting when to add more host machines, etc.
As far as how it compares to k8s, I don't know, I haven't used it materially yet.
Nomad is trying to be an orchestrator.
Kubernetes is trying to be an operating system for cloud environments.
since they aim for being different things they make different trade offs.
How do you scale it, how do you manage it, how will it get deployed, all questions that go into the answer of what should go into it.
Containerfile vs Dockerfile - Infra as code
podman vs docker - https://podman.io
podman desktop companion (author here) vs docker desktop ui - https://iongion.github.io/podman-desktop-companion
podman-compose vs docker-compose = there should be no vs here, docker-compose itself can use podman socket for connection OOB as APIs are compatible, but an alternative worth exploring nevertheless.
Things are improving at a very fast pace, the aim is to go way beyond parity, give it a chance, you might enjoy it. There is continuous active work that is enabling real choice and choice is always good, pushing everyone up.
Afaik, the kernel network APIs are pretty complicated so it's fairly difficult to expose to unprivileged users safely
When I changed my setup from expensive Mac Books to an expensive work station with a cheap laptop as front end to work remotely this was the best configuration I found.
It took me few hours to have everything running but I love it now. New project is creating a new container add a rule to iptables and I have it ready in few seconds.
Exposing the docker daemon on the network and setting DOCKER_HOST I’m able to use the remote machine as if it was local.
It’s hugely beneficial, I’ve considered making mini buildfarms that load balance this connection in a deterministic way.
If you have a secure network then it’s perfectly fine to expose the docker port on the network in plaintext without authentication.
Otherwise you can use Port forwarding over SSH.
To set up networked docker you can follow this: https://docs.docker.com/engine/security/protect-access/
I’m on the phone so can’t give a detailed guide.
BTW, no need to expose DOCKER_HOST, you can connect to docker over ssh, e.g. `DOCKER_HOST=ssh://1.2.3.4`.
Based on my limited interactions with it, I'd recommend staying away from LXC unless absolutely neccesary.
When you run lxd init there's an option to make the server available over the network (default: No), if enabled you can host images from there.
lxc remote add myimageserver images.bamboozled.com
lxc publish myimage
lxc image copy myimage myimageserver
lxc launch myimageserver:myimageIn the end I just settled from hosting on s3, downloading the images with curl (or similar) and running `lxc import`.
The support in configuration management tooling is also pretty limited and the stuff I've used has been fairly buggy, in my opinion, because it's not very popular.
These issues can be observed in the official Ubuntu image and seem to get worse over time. I would recommend to just use VMs instead.
Combined with ipvlan I can flexibly assign my dedicated server’s IP addresses to containers as required (MAC addresses were locked for a long time). Like, the real IP addresses. No 1:1 NAT. Super useful also for deploying Jitsi and the like.
I still use Docker for things that come packaged as Docker images.
Friend, do you have documentation for this process? Please share your knowledge. ^_^
The significant changes from the physical systems were:
* rc_provide="net" in rc.conf because base networking is controlled externally
* rc_sys="lxc" may or may not be necessary
* Disable various net setup services
On the host OS (Debian) I have interfaces like this:
auto ipvl-main
iface ipvl-main inet manual
pre-up ip link add link eth0 name ipvl-main type ipvlan mode l2
post-down ip link delete ipvl-main
In the container config, they are referenced this way: lxc.net.2.type = phys
lxc.net.2.link = ipvl-main
lxc.net.2.ipv4.address = 1.2.3.4/29
lxc.net.2.ipv4.gateway = 1.2.3.1
lxc.net.2.ipv6.address = abcd::2/128
lxc.net.2.ipv6.gateway = fe80::1
lxc.net.2.flags = up
Later on, I removed the dedicated IP address and set up a reverse proxy instead.Oh yeah, all containers are of course privileged containers. With unprivileged containers, various things may not work as expected.
I couldn't have said it better. And yes, I use it. Also in production systems.
I haven't seen many examples demonstrating the tooling used to manage LXC containers, but I haven't looked for it either. Docker is everywhere.
But Docker is many things: a company, a command line tool, a container runtime, a container engine, an image format, a registry...
[1] https://sarusso.github.io/blog_container_engines_runtimes_or...
That's all I've ever needed. Docker is overkill if you just need to run a few containers. There is a point where it makes sense but running a few containers for a small/personal project is not it.
In my mind you need to treat LXC containers as VMs, in terms of managements. They need to be patched, monitored and maintained the same ways as a VM. Docker containers still need to patched of cause, many seems to forget that bit, but generally they seem easier to deal with. Of cause that depends on what has been stuffed into the container image.
LXC is underrated though, for small projects and business it can be a great alternative to VM platforms.
LXC is a fantastic userland library to easily consume kernel features for containerization without all the noise around it… but the push for the LXD scaffolding around it missed the mark. It should’ve just been a great library and that’s how we use it when running containers on embedded Linux equipment
[1] https://pantacor.com/blog/lxc-vs-docker-what-do-you-need-for...
And still having troubles running most of the docker images out there (either this, or that won't be supported). I guess it makes sense, after all there is always the choice of going with full real linux reinstall, or some other hacky ways.
But one thing I was not aware was this: "Docker containers are made to run a single process per container."
There are a plenty of other solutions and Docker is actually many things.. You can use Docker to run containers using Kata for example, which is a runtime providing full HW virtualisation.
I wrote something similar, yet much less in detail on Docker and LXC and more as a bird-eye overview to clarify terminology, here: https://sarusso.github.io/blog_container_engines_runtimes_or...
“ LXC, is a serious contender to virtual machines. So, if you are developing a Linux application or working with servers, and need a real Linux environment, LXC should be your go-to.
Docker is a complete solution to distribute applications and is particularly loved by developers. Docker solved the local developer configuration tantrum and became a key component in the CI/CD pipeline because it provides isolation between the workload and reproducible environment.”
LXC on the other hand is lightweight virtualization and one would have a hard time to use it without basic knowledge of administering Linux.
So what is Docker doing then??
I tried docker but stuck with lxc.