The Case for the /usr Merge (2012)
freedesktop.org
freedesktop.org
> The primary commercial Unix implementation is nowadays Oracle Solaris.
Surely it's macOS? The total install base of Solaris must be a rounding error compared to the number of Macbooks sold daily.Previous recent discussion on that specific quote: https://news.ycombinator.com/item?id=30929464
In some senses macOS is more of a UNIX than Solaris is, as it is certified by The Open Group and Solaris is currently not: https://www.opengroup.org/openbrand/register/
I think a lot of people who think "macOS isn't Unix" (or similarly, "z/OS isn't Unix") aren't aware of the legal status of the term "Unix" – that it is a registered trademark owned by The Open Group, and legally a trademark means whatever the trademark owner says it means – and The Open Group is very clear that "Unix" means any system, irrespective of its marketing or heritage or other functions, which passes their test suites and whose vendor pays the licensing fee.
The phrase means Sun Microsystems, HP, IBM and SGI variants of Unix.
And does “IBM… variants of Unix” include z/OS? If not, why not?
HN thread on this: https://news.ycombinator.com/item?id=29984016
And for a couple of years after they sorted out the compliance, they really pushed UNIX in their advertising materials. After all, they paid more than $20mn to sort out compliance.
UNIX is a certification. macOS is certified. Solaris is not and hasn't been since 2020.
I am not sure what you’re implying. There always have been a lot of UNIXes that were “commercial OS”. Historically, the free branches of the UNIX family tree have been way out-numbered by commercial distributions.
> but does not tend to show it too much
That does not really matter (besides the fact that the terminal has always been there). The fact that there is a GUI shell on top does not change the architecture of the OS.
> The Unix part is an implementation detail.
These days, UNIX is a certification. Otherwise it refers to the architecture and origins of the OS, from the initial AT&T UNIX. The only purity test is the certification.
> It is not marketed as a Unix workstation, much like Android phones are not marketed as Linux handhelds.
How it is marketed and what it actually is are orthogonal concepts. It is UNIX, whether this annoys some people or not.
(importantly those links still work, even outside of the Wayback machine)
Solaris, BSD, Darwin all share code.
Linux is POSIX compatible and userspace is similar, but the kernel is distinct.
Why spend all that money and time when Linux has so dominated the certified Unices for some time?
Looked it up, here's the page for the OS: https://developer.huaweicloud.com/ict/en/site-euleros/eulero...
And here's a news post about the certification: https://www.huawei.com/en/news/2016/9/huawei-kunlun-euleros-...
> [Shanghai, China, September 9, 2016] Today Huawei announced that its EulerOS 2.0 operating system for its KunLun mission critical servers has earned UNIX 03 certification from The Open Group. This means that Huawei EulerOS 2.0 has been officially certified as a UNIX system.
Inspur K-UX, on the other hand, is a Linux distribution which is a Unix
https://en.wikipedia.org/wiki/Inspur_K-UX
> Inspur K-UX 2.0 and 3.0 for x86-64 were officially certified as UNIX systems by The Open Group.
[snip]
> Inspur K-UX 2.0 was one of six commercial operating systems that have versions certified to The Open Group UNIX 03 standard (the others being macOS, Solaris, IBM AIX, HP-UX, and EulerOS).
More:
https://en.wikipedia.org/wiki/Unix
> Systems that have been licensed to use the UNIX trademark include AIX,[35] EulerOS,[36] HP-UX,[37] Inspur K-UX,[38] IRIX,[39] macOS,[40] Solaris,[41] Tru64 UNIX (formerly "Digital UNIX", or OSF/1),[42] and z/OS.[43] Notably, EulerOS and Inspur K-UX are Linux distributions certified as UNIX 03 compliant.[44][45]
None of the Open Source BSDs are on that list.
EDIT: To make it more explicit, given that people keep mistaking Android for Linux,
"NDK C library"
https://developer.android.com/ndk/guides/stable_apis#c_libra...
"Improving Stability with Private C/C++ Symbol Restrictions in Android N"
https://android-developers.googleblog.com/2016/06/improving-...
"Android changes for NDK developers"
https://android-developers.googleblog.com/2016/06/android-ch...
Google for phrases like "idg unix shipments solaris" and you'll see some of the context from that time period.
Back in 2005 or 2006, I made a Linux that exported the filesystem this way, and it worked by mounting / as a network volume. This worked by doing a particular dance with "pivot_root" at the beginning and a union mount with a tmpfs volume. We used it to play LAN games in the computer lab. Because the network booting Linux didn't touch the local hard drive, you could clean up after yourself just by shutting down the machines. It was also nice and fast, because the machines all had gigabit ethernet and the file server was serving from cache most of the time. It ended up being noticeably better than running off local hard disk... this was the era before SSDs took over, obviously.
The main difference is about static linking.
> /bin/ User utilities fundamental to both single and multi-user environments. These programs are statically compiled and therefore do not depend on any system libraries to run.
> /sbin/ System programs and administration utilities fundamental to both single and multi-user environments. These programs are statically compiled and therefore do not depend on any system libraries to run.
It was discussed a few month ago: https://news.ycombinator.com/item?id=31336396
Basically, glibc uses dlopen to open certain shared libraries for things like nsswitch and locales. Using dlopen in a program that's statically linked to glibc causes all kinds of issues, as you'll likely end up with multiple copies of potentially different versions of glibc in the same process.
There is some case to throw /usr/share on separate partition as some packages dump quite a lot data there (kicad for example dumps part libraries there which are like 5+GB), but that's the single case in "basically never" I had something mounted in or under /usr.
But honestly, if you have a case of "I REALLY need a rescue environment that's statically compiled".... just throw a busybox binary on your root partition somewhere. Most distros have it in "busybox-static" package.
> It was discussed a few month ago: https://news.ycombinator.com/item?id=31336396
That just seems mostly BSD being silly stubborn to be honest.
No? The distinction could also be self-imposed, i.e. you just put statically linked "rescue" binaries in /sbin just for the sake of order/principle, even if it happens to be in the same disk or the same partition as /usr.
It may be absolutely pointless, but then so would be literally any other partitioning of files you can think about. E.g. what's the point of creating a directory for each program ala Windows/OSX? What's the point of making sbin vs bin? What's the point of putting data files in share instead of lib? What's the point of creating any hierarchy whatsoever instead of just throwing everything inside the same directory and let the package manager handle it?
You can find very thin arguments for all of these, most of them having to do with historical reasons or subjective reasons, like some human abstractly-defined sense of simplicity, "cleanliness" and order.
- it separates /, /usr and /usr/local
- it has special mount options such as wxallowed which is usually enabled on /usr/local only
https://man.openbsd.org/mount.8 / https://man.openbsd.org/fstab.5
> /sbin/ System programs and administration utilities fundamental to both single and multi-user environments. Most of these programs are statically compiled and therefore do not depend on any system libraries to run.
I didn't follow the news, thanks for the update :)
Though I responded in 2012, so it may be older (the log on the site says 2013 though).
FWIW, my response, as FHS editor:
So I can kinda understand the blind hate
[1] https://github.com/systemd/systemd/issues/5644#issuecomment-...
You can just smell man have attachment issues with his design decisions and code. Reimplementing stuff that works fine, but badly also seems to be his fucking hobby
(Assume someone was willing to write and supply the text.)
I don't know what this guy is smoking, but this has never been a problem.
Fix that one, and you can use
#!/usr/bin/env interpreter --arg --arg ...
Problem solved; no need to mess with the file system divisions.For that matter ... the stupid hash bang mechanism itself should do PATH resolution, maybe, you know?
#!/usr/bin/awk -f <-- find awk at that absolute path
#!awk -f <-- Just use the darn PATH, dear hash-bang handling code in the kernelNot to mention library linking and looking for various other data files.
I've run into this particular issue with stripped down containers.
I run a script as part of the container's CMD/ENTRYPOINT, and it fails.
Run it through an interactive shell, and hey-presto PATH is setup and script runs fine.
Perhaps if you started your entrypoint with “/bin/bash cmd”??
#!/usr/bin/env -S interpreter --arg --arg
https://www.gnu.org/software/coreutils/manual/html_node/env-...Here is a problem: interpreters take arguments such as options ... but, separately, so do interpreted programs. Sometimes, you'd like to control both.
I made an extension to env whereby you could do:
#!/usr/bin/env :interp:--foo:{}:--bar
This would find the "interp" interpreter using a PATH search and then pass it the arguments --foo <scriptname> --bar, followed by arguments that came from the invocation of the hash-bang script. If the special argument {} does not appear, then interp will be invoked with --foo --bar <scriptname>: the script name is not relocated into the arguments.However, the maintainer expressed a preference for compatibility with FreeBSD -S.
I took a look at the requirements and saw that FreeBSD's env -S was doing a whole lot of stuff: interpreting C-like character escape sequences like \n, and interpolating dollar-sign-sigiled variables like $USER. In spite of all that, didn't see the feature of being able to insert the script name into the middle of the arguments, like in my solution.
It amounted to throwing away my patch and implementing some FreeBSD stuff I had no interest in and a lot of which I thought was a bad idea, while failing to achieve my original functionality. So, I just did the first part: throwing away my patch.
But why would a script ever need to use a shebang to pass an argument like --bar to itself? Can’t it just modify its own argv, or act as if it had?
(I found your message at https://lists.gnu.org/archive/html/coreutils/2017-05/msg0002... that claims
> It is useful because it allows the hash bang to specify some arguments after the script (which could be arguments belonging to the script rather than to the interpreter, for instance).
but doesn’t really explain why the script itself would want to specify that.)
In any case, compatibility is important here, in order for it to be possible to write cross-platform scripts.
Suppose that the interpreter is such that the script name must be passed as an option. For instance, Awk implementations are like this:
awk -f <script>
The problem is that this kind of option is allowed to be followed by more options. awk -f <script> -Z # error under GNU awk: -Z is an invalid option
See where this is headed? If we have #!/usr/bin/awk -f
awk script here
and we call this as $ ./script -Z
it will pass that option to Awk. But the script wanted to handle that option! $ ./script -- -Z # process -Z as argument to the script
and wouldn't it be nice if we could hide this "--" argument in the hash bang?With my proposal you could do that:
#!/usr/bin/env :awk:-f:{}:--
Now speaking of Awk, GNU awk has solved this particular problem itself. It has the option -E/--exec which works like -f, but is the last option to be processed.(Any of these kinds of issues can be solved locally for a given interpreter, if you control its implementation.)
I wrote [0] in lua a while back (mind the comments), didn't destroy anything I think, I wrote [1] just now in gawk, so not as recommended.
FWIW the correct solution, is the plan9 one, just union mount/bind all paths to /bin.
-> ᛯ cat /tmp/1.sh
#!echo ble bla bla bla -n -e
-> ᛯ /tmp/1.sh
ble bla bla bla -n -e /tmp/1.sh
It's the best solution but it will take decades for that to migrate to other shellsThe exec fails when a shell tries to run an executable file that has no hash bang header, but also in a case like this when the header doesn't give a path that resolves to an interpreter.
Without the #!echo, I'm guessing Zsh would have interpreted the file as as Zsh script.
It's really not a huge issue for me, but it is at least a minor frustration. (And for people who maintain scripts across different distros / *nixes, I'm sure that it is a larger frustration).
I still write env bash shebangs to this day, in his honor. Happy to hear there's progress towards ending this archaic separation.
POSIX explicitly says "Applications should note that the standard PATH to the shell cannot be assumed to be either /bin/sh or /usr/bin/sh, and should be determined by interrogation of the PATH returned by getconf PATH [...]"
Worse: it’s in /usr/local/bin now (as ports are installed in /usr/local, by archaic convention).
To assume otherwise would have to assume that all ports come from FreeBSD and that's a bit absurd. I don't think anyone's going to say that Bash, or say KDE, is a "FreeBSD project" any time soon.
And all that just to point out that /usr/local on FreeBSD is intended for non-OS software installs (it's not the only option you have on the system, /opt isn't unheard of, but mixing up /usr/bin with non-OS software is a fast recipe for disaster).
It does make sense for your pythons and rubies but not really for bash.
Side note: zsh allows for #!bash but that's only interpreter I found that does that
<snort>
I'd say the biggest issue is that the one benefit I'm finally reaping, and apparently mentioned (badly) at the time, is the one benefit that all the mentions of merging explicitly disavowed as impossible and that dinosaurs like me should go and Bury themselves.
Namely, separate /usr partition made readonly. Serving readonly FS or shared FS for systems has been harder with unmerged /usr because of /etc (among others).
However, the arguments for merging that I heard were "nobody does separate /usr anymore you dinosaur", "nobody wants to share filesystem between multiple system instances", and finally from the proposers themselves, "split / and /usr mean I have to keep track of where I'm putting files in RPM depending on whether they are early boot or not, boohoo".
So recently I've been working on building a customised Gentoo, based on ChromeOS/CoreOS/Flatcar changes which tells emerge to use /usr as target... And for the first time in forever I have use for merged /usr and it's an use that was ridiculed when merged /usr was proposed (but use that was explicitly (ab)used by Solaris, btw)
/usr/bin -> /bin
One less path component to resolve.
Fact: This would make the separation between vendor-supplied OS resources and machine-specific even worse, thus making OS snapshots and network/container sharing of it much harder and non-atomic, and clutter the root file system with a multitude of new directories."
In general, the desire is to put the system image into /usr. A single partition, that holds literally everything you need to launch an os. If you can mount a /usr, everything else will populate. If you want to atomically swap system images, just swap /usr.
Imagine the counter: what would happen if we just swapped /bin? Shit would be a mess. /lib and /usr/lib wouldnt get updated in sync. /include and /share wouldnt be in sync. The idea is that the Unified System Resources ought unifiedly be swapped out. And the system specific resources(/bin, /lib, /include, et ceteta) ought not exist, provide no real contemprary value, and making image swapping a much much harder to accomplish well.
FYI “usr” really means “user”. Somebody made up “unified system resources” as a backronym.
A lot of it relies on better filesystems like ZFS or Btrfs which have snapshots. I myself twiddled around with writing some Debian-Live boot tools that used btrfs snapshots, copied onto ram-drives[2], that would potentially work well with package managers (boot either your starting image, or your "new" snapshot image).
These ideas predate (the deeply reviled by some haters) Lennart's post, "Revisiting How We Put Together Linux Systems"[3]. But Lennart here, in mid 2014, understood what containers were doing, understood how atomic system upgrades were trending, understood what was happening. This isn't a Lennart idea, but Lennart as usual semi got the gist of what made sense, what would help, where we were going. He captured the good idea of what was occuring already into a post. And calling out /usr as "the system image" was a clear & obvious win for that day, back in 2014, and it's still a clear & obvious win today to consolidate the OS into one directory today. May it be so. This is the way. So say we all. It is known.
The "usr" meaning user is seemingly in fact true. But it's also, imho, wrong? The backronym is more truthful, better describes what role /usr plays & has played for multiple decades now. "user" doesn't really mean anything reasonable, clear, or cogent.
[1] https://gitlab.com/wyrcan/wyrcan
[2] https://github.com/rektide/debian-live-boot
[3] https://0pointer.net/blog/revisiting-how-we-put-together-lin...
They do some funky stuff with /usr/local and a few other directories, too. This person seems to explain it: http://www.swiftforensics.com/2019/10/macos-1015-volumes-fir...
Deployed desktop Linux has /bin and /lib, so it's somehow working.
- RO /usr contains all the system, the rest is populated with tmpfilesd
- There is two /usr partitions, the active one is selected with GPT priority attribute
https://github.com/endocode/coreos-docs/blob/master/os/sdk-d...
Hasn't this use case been replaced by initramfs?
So the common argument was that one should accept that readonly separate /usr and snapshots and whatsoever you wanted weren't going to be a thing because fedora didn't want to maintain support for mounting /usr in initramfs.
It looks like a bad idea that I would never want to do.
This is what your chroots and jails are for.
But if I wanted to do it in style, I'd hack up the cool kernel support to atomically mount multiple directories as part of one atomic swap.
How can you call it "atomic" if you have executables still running from the old /usr/bin? Say there is some script running that is executing external programs. At some unspecified point in its execution, /usr/bin/whatever changes meaning; all newly run programs are different versions on a different libc and so on.
If you have enough control of the system that there are no such processes, then you can easily do the swap directory by directory: just do it with system calls out of a single process, rather than by invoking tools, so that nothing trips up at the point in time when you have a mismatched /bin and /lib.
These upgrades often involve a kexec restart. For fleets of servers- the target audience- it's expected that machines will go down & rolling restarts should be fine. That's how true atomic behavior is gotten. But I think you are nitpicking & being fussy with your complaining about old processes being left around; even if you do an in-place swap, but have some old processes around, at least the whole system is together & consistent. More broadly, thjs set of problem in general doesnt change much whether you are using a merged /usr or not.
I genuinely dont get any of your disgust at having a /usr image that contains the functioning system image. You talk about chroots, as though it's an alternative somehow, but this practice of having a self-contained image seems like it would greatly assist in quickly cloning new images to chroot into.
It should make a lot of tasks a lot easier. The one I like most is trivially setting up containers or VMs that have your base OS in them plus some new software that you are testing. For example, at work we use Debian with a bunch of extra packages installed, plus our own package containing all manner of software and configuration that we’ve written. Although we do a fair amount of unit testing, and the Rust compiler really has our back, it is very difficult to really test things outside of production. Currently we have to painstakingly set up a test environment on a production machine, using tricks like altering our PATH so that the test version of the tools we are working on will be called, or editing things so that they call the version under test instead. It’s doable but laborious.
If we were using Fedora instead of Debian, we would be able to use systemd-nspawn to spawn a container with the existing OS, read–only mounts containing our databases and other large datasets, and an overlayfs to redirect any writes to those file systems to an alternate location. Installing and running our software inside that container would work exactly as if we were in production, and testing everything would be so much simpler.
systemd-nspawn is available on Debian, but running it requires keeping extra copies of the OS for each container you want to spawn, so we haven’t actually gotten around to doing it yet. Oh, and each container needs a properly set up /etc directory. We can’t just leave it empty, and we can’t just blindly copy the host’s either.
A common example of "what's wrong with split /usr" that was pushed was things like dropping pci.ids database in /usr/share and it being required by code mounting /usr, and that mounting / and /usr separately in initrd was a wasteful complexity.
In fact, today is the first time I have heard of reasoning that supported /usr on separate partition and I thought it was just an inversion done later by people who noticed possible benefits for some extra initrd complexity (btw, dracut and systemd oriented initrd scripts are stupidly complex. Including generating systemd units by text templating in udev rules as part of dracut boot process. Wtf)
The Unix way will prevail. I am happy the BSDs continue with the philosophy. Old does not mean outdated and tradition does not mean backwards. One thing only and do it well.
> Improved compatibility with other Unixes/Linuxes in behavior:
Well, some of them.
> Improved compatibility with other Unixes (in particular Solaris) in appearance:
I guess, but if you supported Linux you already supported Solaris, so if you think of Solaris as second class then nothing gained.
Also: Solaris is (to me) clearly on its way out. Not worth a migration to converge with it.
> Improved compatibility with GNU build systems:
This one is pure hogwash. Either you need to support the systems that don't do this (e.g. OpenBSD), or you only care about Linux. (this is a bit simplified, but pretty true)
So your build system still needs to support non-merged.
Except now you're going to bake in Linux-specific merge assumptions into your build system, so that they'll be even more broken.
> Improved compatibility with current upstream development:
This seems more like tech debt to me, in the name of expedience.
I'm here not saying it shouldn't be done. I'm saying these are terrible reasons.
> This one is pure hogwash. Either you need to support the systems that don't do this (e.g. OpenBSD), or you only care about Linux. (this is a bit simplified, but pretty true)
> So your build system still needs to support non-merged.
> Except now you're going to bake in Linux-specific merge assumptions into your build system, so that they'll be even more broken.
Well considering
> Not implementing the /usr merge in your distribution will isolate it from upstream development. It will make porting of packages needlessly difficult, because packagers need to split up installed files into multiple directories and hard code different locations for tools; both will cause unnecessary incompatibilities. Several Linux distributions are agreeing with the benefits of the /usr merge and are already in the process to implement the /usr merge. This means that upstream projects will adapt quickly to the change, those making portability to your distribution harder.
I think they intended this to play out like SystemD/logind - Red Had maintainers will only care about merged /usr and if you want to use anything they have their hands in you better adapt.
The argument this article seems to make is that you can essentially hard code the paths to binaries in your build system, because /lib is /usr/lib, /bin is /usr/bin (from build script point of view it doesn't matter which is symlink to which).
It seems to be saying that your build systems no longer need to search for the libraries.
But that's not true. You still do.
And if you search for it then it doesn't really matter where it is. Clearly if you search for `bash` you'll search /bin and /usr/bin. For at least two reasons:
1) Maybe you're running a system that isn't merged. Then it could be in either place. Doesn't matter if RedHat is merged.
2) You'll need to check other locations anyway. Like /usr/local/bin/bash, to support some BSDs. Why does it matter if it's two, three, or $(echo $PATH | sed 's/[^:]//g' | wc -c) paths?
And for libraries you may need to search in /opt, and run pkgconfig for any compile and link flags needed. Why does it matter that `pkg-config --libs foo` contains an -L flag? You still need to run it.
I think even /usr/share can get its own directory, why not?
Because then you can mount /usr from wherever, which is much harder to do with /. thus merging in /usr provides utility which are harder and less convenient to provide the other way around.
tl;dr:
/usr/local - Place for the system administrator to put local stuff. That is, the package manager (apt/dnf/yum/whatever) won't touch anything there.
/usr/bin - "Normal" binaries
/usr/sbin - System binaries, that is, binaries that don't do anything useful when run without superuser privileges
/bin - Well, nowadays just a symlink to /usr/bin. Read the article to understand the historical background why they were split.
/opt - For packages that install everything under one directory instead of being spread out into the standard locations.
These days people just invent some fiction about how it’s actually “Unix system resources” and not just a second user folder.
1. /usr/local is usually set up to take priority over /usr, so if you want a newer C compiler or whatever you put it in /usr/local but keep the system compiler in /usr because you might need it for an OS patch or something.
2. If you see an error from something in /usr/local you know not to bother your OS vendor about it because it came from an outside source.
It's largely the same theory as /opt, except stuff in /opt shouldn't link to any libraries in /usr while stuff in /usr/local can (stuff in /opt should be self-contained).
https://lwn.net/Articles/773342/
tl;dr: Do the merge or don't but don't support both.
make /usr home again.
Since when are you so concerned about Unix compatibility as you push systemd?
- If you have your own system, with very tightly controlled specifications, you can make any change you want, and its impact will be easy to manage. When you change the specifications of random people's systems, random things happen. Doesn't matter what the change is. A file showing up where it didn't exist before, or being a symbolic link rather than a regular file, or duplicate files, will cause logic bugs.
- There isn't a single mention of any downside to this fundamental shift in the expectations of the filesystem of major distributions. Yet you can be guaranteed that this will break compatibility with some applications/systems. They still they make no attempt whatsoever to estimate this.
- It can't be stated enough that this is a Fedora / RedHat / Lennart Poettering thing. Quite frankly, anything they suggest should ring giant alarm bells. This group is singly responsible for every incompatible, over-designed, pain-in-the-ass change to Linux distributions in the past 10+ years. If it causes problems, and it didn't come from a hardware vendor, it came from these people.
- Who in the actual fuck cares about filesystem-level porting from Solaris?! If you're already dealing with all the rest of the portability issues, file paths are literally the very last and least-difficult consideration. "Oh no! In what directory will I find 'pkill' ?! It's so difficult!" Only a company obsessed with converting whales to their platform and want one more bullet point to add to their sales deck cares. I don't actually care if this goes through, but just the fact that RedHat is still shilling their unnecessary changes as if anyone else other than them wants it makes my skin scrawl.