FreeBSD Jails for Fun and Profit (2020)
topikettunen.com
topikettunen.com
Here's an example from my personal name server:
/usr/sbin/jail /jails/www www 10.10.10.36 /lighttpd -f conf/lighttpd.conf
... and although this jail has a lot of content files in it, the actual UNIX userland is only what is required to run 'lighttpd': # find /jails/www/usr | wc -l
43
So it's an extremely lightweight environment with very little attack surface.You can also share a lightweight environment with multiple commands - here are two other jail commands:
/usr/sbin/jail /jails/dns ns1 10.10.10.30 /nsd/nsd -c /nsd/nsd.conf
/usr/sbin/jail /jails/dns dns 10.10.10.37 /unbound/unbound -c /unbound/unbound.conf
... see how both jailings of 'nsd' and 'unbound' point to the same '/jails/dns' userland ? Once again, that userland is very, very compact: # find /jails/dns/|wc -l
97
... so, 97 files total to run both name servers.No 'make world' necessary, no building and maintaining of a full FreeBSD system - just the lightest skeleton required for both 'nsd' and 'unbound'.
Packaging an entire system is more about convenience than anything else. It's also pretty difficult to package just the libs one needs when you are dependent on libc and other C libs.
I suspect that if one was really ok with it, some tooling could be built to copy/link in system libs into the rootfs automatically from the host.
It's pretty easy using Nix, e.g. this example defines a container for running a shell script: https://ryantm.github.io/nixpkgs/builders/images/ocitools
That script depends on bash, bash depends on libc, etc. so those dependencies (and only those dependencies) will be put in the container. (See https://nixos.org/guides/nix-pills/enter-environment.html#id... for an example of what dependencies look like in Nix).
> I suspect that if one was really ok with it, some tooling could be built to copy/link in system libs into the rootfs automatically from the host.
Eww, no thanks! I want my containers to be reproducible.
Exodus – relocation of Linux binaries–and all of their deps–without containers - https://github.com/intoli/exodus
I did some reading and it looks like FreeBSD's unionfs or nullfs would handle the file-system part.
Just seems unnecessarily complicated.
Going further, you can place the few executable files you need in a read-only filesystem, so attackers can't copy their own payloads in, further restricting what they can do. You can continue that process with firewall rules that are much more restrictive than what you could use with a more general purpose server, such as blocking any traffic that isn't to or from port 80/443. You can also virtualize the network of a jail, so an attacker wouldn't even get information about the network's layout from a compromised jail.
Typical application attack path:
RCE in the app -> ability to use anything in the OS under the app user privileges -> privilege escalation -> The whole box pwned (including access to all your secrets and management credentials)
Remediation process: wipe the box
Attack path for an app within a jail:
RCE in the app -> ability to use anything in the OS but there is nothing to use, must bring own tools (not always feasible) -> privilege escalation -> JAIL is pwned but not the box
Remediation process: recreate the jail
[...] In fact, many years ago, when FreeBSD was my main OS (including on notebook) I went as far as to isolate each app that used internet into its own custom-setup jail [0][1]. I had Firefox, Thunderbird, Pidgin and a few others running in complete isolation from the base system, and from each other. I even had a separate Firefox jail that was only allowed to get out via a Tor socks proxy to avoid leaks (more of an experiment than a necessity, to be fair). Communication between jails was done via commonly mounted nullfs. I have also setup QoS via PF for each of them. They were all running on the host’s Xorg, which was probably also the weakness of this setup. It was a pretty sweet setup, but required quite a bit of effort to maintain, even tho I automated most of the stuff. [...]
The original comment is here: https://news.ycombinator.com/item?id=27709256
Yes, there's always a question of usability, but if you're an advanced user, it's still scary that you're not afforded any control to prevent these incidents unless you go ahead and redesign the whole way all these apps are working all by yourself.
It seems terribly inefficient if every engineer has to do it on their own, and in their own incompatible way. Obviously most people simply give up after a while, since maintaining such a setup might itself be a whole full-time job.
I have looked into doing this many times and it's neither simple nor straightforward.
Specifically: jailing a GUI app that you can interact with on your desktop.
I can't remember what the most promising recipe I saw for this was but it wasn't quite promising enough to compel me to built it up ... and this discussion is always (rightly) hijacked with "just use Qubes" ...
There's also a very interesting read on Qubes-like experience on NixOs with Wayland and XWayland[2,3].
[1]: https://wiki.gentoo.org/wiki/User:Sakaki/Sakaki%27s_EFI_Inst...
[2]: https://roscidus.com/blog/blog/2021/03/07/qubes-lite-with-kv...
In my case, I have the standard xorg session started by my login manager. Then I start Xephyr with a separate DISPLAY, that shows up as just a window in the parent environment. It does look kinda like RDP or VNC.
[1]: https://linux.die.net/man/1/xephyr [2]: https://wiki.archlinux.org/title/Xephyr
I wasn't aware of HP-UX Virtual Vaults, thanks. However, I'd say FreeBSD Jails still had an advantage, due to being free and running on various hardware platforms, and more importantly on commodity hardware.
Unfortunately it is hard to still find documentation, given the troubles HP-UX has gone through at HP (which kind of plays into your remark regarding FreeBSD).
Still there you go, https://support.hpe.com/hpesc/public/docDisplay?docLocale=en...
What the market actually wanted was Docker. And what Docker needed was Linux containers (complicated, flexible, piecewise technology) and not jails, which were higher level abstractions (but yet not high enough) with jargon and framework assumptions that didn't match Docker's needs 1:1.
There are many reasons why FreeBSD jails count not get out for FreeBSD land, one, very important thing is the Linux community's NIH attitude.
- brtfs vs zfs
- cgroups vs jails
- SystemTap vs dtrace
- Systemd vs smf
I get it, many of these were due to licensing issues. So they said[1]. Anyways, there are still some things to implement for linux. pf is my favourite (software) firewall. It would be great to see it ported to Linux.
1. https://opensource.stackexchange.com/questions/2094/are-cddl...
BSD jails are similar but not quite the same thing.
One real difference is that you need to be root to create a jail. It'll get fixed eventually - FreeBSD already has unprivileged chroot, jail isn't that much different.
Not really, the example of Docker would probably be the most straightforward there. I don't think it's possible to fully port Docker to jails or at least I've never seen a successful port, some of the network topology features seem to just not be possible or straightforward. But I could be wrong, I have not looked into the technical details of this in years, somebody told me it might have been working a while ago but I never heard anything else about it since.
Needing to be root is a major deficiency though and I can't take jails seriously with that, one of the main focuses on Linux containers in the past several years has been to make unprivileged namespaces a good option.
And yes, having to use root is a major issue. Looks fixable though.
>And yes, having to use root is a major issue. Looks fixable though.
AFAIK it took a long time to get this to work on Linux, there are a lot of security issues that it can cause.
Many things are hellishly complicated in Linux, due to politics and technical difficulties. Case in point: when I’ve started to work on NFSv4 ACLs, support in Linux was “worked on”, there was a prototype. It was 12 years ago. In FreeBSD, full supper for NFSv4 ACLs, from file systems to userspace tools, shipped decade ago. In Linux it’s still not there.
BSD doesn't really have any licensing issues, thanks to BSD license, but politics is directly related to project size. In FreeBSD it's pretty much unnoticeable, but in Linux it can be a huge deal.
FreeBSD does avoid pulling restrictively licensed (closed source or GPL) code into the base system itself, but a Docker port would be third party (ports/packages), not the base system.
Also, it’s not politics - it’s mostly just that the old GNU cruft is being replaced, and newer, better solutions prefer more liberal licensing, see GCC vs LLVM.
You are right, but ironically we have to thank Microsoft for that ;)
Note: Linux also needs root for its namespaces. Or at least CAP_SYS_SYSADMIN, which grants enough that it's pretty much as good as root. See setns(2) and clone(2) for details. This is one of the complaints the plan 9 people have always had with Linux namespaces.
The sandboxing and mount-related ones are implemented with namespaces, and the idea with them is to not make any of them mandatory so they can be slowly added to system services. That way you can get some of the benefits without needing to build a full rootfs/container for the service. I am not sure how any of those would be done with jails because jails require you to create a chroot and network interface, whereas in Linux the mount and network namespaces are just optional namespaces and you can still use the other namespaces without using them.
They don't: you may chroot to /, share the host's network interface, or disable networking.
trasz@v3:~ % doas jail / foo 127.0.0.1 /bin/sh
# ps aux
USER PID %CPU %MEM VSZ RSS TT STAT STARTED TIME COMMAND
root 37975 0,0 0,0 13500 3056 3 SJ 09:11 0:00,01 /bin/sh
root 37976 0,0 0,0 13624 2776 3 R+J 09:11 0:00,00 ps auxUm... to loop back to the upthread point: Docker. People are using Docker, and docker is using this stuff.
If I'm wrong: what is it using, and what problems is this flexibility solving?
And runj, IIRC (though I'm not an expert in the space) wasn't a trivial 1:1 thing and required changes to the underlying jails layer to enable it.
Docker as a paid product failed. Docker as a company is failing. "Docker" in the sense I meant (of the software people use to launch containers), is pervasive and dominant. It won. And jails, in comparison, "lost", because jails didn't really do what Docker wanted. And what the market wanted was Docker.
One thing I like about iocage is how easy it is to grant the jail access to the host ZFS datasets.
On a Jails note, I have had issues creating a jail that can do network inspection. I believe this is an issue with network restrictions of the jails subsystem itself. Eg, I could never run nmap or get mac addresses of remote hosts from within a jail.
The tooling is slowly moving in a direction I like, though :)
This is an old post of mine which I happened to find useful. Orchestration of jails moved quite bit forward lately! For example, you can manage your jails quite nicely with containerd today! See great post from Samuel Karp about the topic: https://samuel.karp.dev/blog/2021/05/running-freebsd-jails-w...
Jails, on the other hand, are not a sandboxing mechanism - they are system-level virtualization, like Linux namespaces, but with a simpler interface. You can use it for sandboxing, but it's not what the mechanism fundamentally is.
apt-cache search jail
firejail - sandbox to restrict the application environment
firejail-profiles - profiles for the firejail application sandbox
firetools - Qt frontend for the Firejail application sandbox
A Docker-like solution with a pretty UI could be really useful for pros. For novices, it could mean a less cumbersome security measure than the restrictions we’ve been experiencing since Catalina.
Worked well from the limited testing I have done so far
Opensource Linux-vserver https://en.wikipedia.org/wiki/Linux-VServer was available in 2001. At the same era commercial Virtuozzo was ubiquitous amount hosting providers.
No, I didn't. I was comparing full virtualization to Jails. However, I realize now that there may not even have been a full virtualization solution available back then anyway, at least not for consumer hardware. I'll have to dig a bit on Wikipedia.
I've never used Solaris Containers / Zones, but my understanding is that the implementation was similar to FreeBSD Jails, so I have no reason to believe the performance was different.
Back in the earlier days of containerisation Linux had no options (Linux was pretty late to that particular game) and Solaris wasn’t free. So FreeBSD made a lot of sense.
These days the tooling around Linux is better and there are open source forks of Solaris so FreeBSD might seem like an odd choice for some. However I still think FreeBSD is a rock solid operating system and one that doesn’t get taken as seriously these days as it should do.
No other OS gives me this kind of comfort and stability.
The base FreeBSD system is well thought out, and because it had the ability to, it coalesced into a very coherent system. Man pages are great, everything is where it should be.
* lower runtime overhead
* quicker deployment time
* smaller image foot print
Plus most of the advantages of VMs can still be applied.
However it does massively depend on individual use cases. Qubes aims to be ultra secure and containers in Linux weren’t up to that task at the time (the situation has since improved massively).
- quicker deployment time
- smaller image foot print
Firecracker beg to differ. I think you can do virtualisation and have all three you listed.
Also I’m not really in the mood to engage in a dumb flame war. I’m just answering the question as to why some people favour containers for some workflows. If firecracker works for you then keep at it.
I tend to look at it the other way though: given containerisation is so easy these days, what are the compelling reasons to run virtual machines.
Neither is a wrong answer though.
The kind of domains where this is an issue isn’t what you’re average engineer would be concerned with.
Where VMs do help the average engineer is they have more secure defaults. It’s pretty easy to accidentally run containers insecurely on Linux (FreeBSD jails are a better story though)
Tells me you don't know the difference between HW-Virtualisation and OS-Virtualisation.
> On your phone, a desktop a server?
Phone doesn't support hardware virtualization of course. On my laptop I'm using Qubes OS.
Every CPU "supports" hardware virtualization.
>software virtualization?
That's a "VM" like JVM, we talk about OS-Virtualisation.
For instance, awk can act differently and if you don't know the 25 year old decisions that led to the differences, it can be very confusing. (There are different behaviors and command line switches between GNU's AWK and the version from SVR4/XPG)
You can't even see a binary from the rest of your system, and exec won't get you out.
[0]: https://docs.microsoft.com/en-us/virtualization/windowsconta... [1]: https://docs.microsoft.com/en-us/windows-hardware/design/dev...
https://docs.virtuozzo.com/virtuozzo_hybrid_server_7_users_g...