Cosmopolitan Libc 1.0
github.com
github.com
It looks like a very clever way of packaging an x86 fat binary for multiple platforms, without actually duplicating the code. To support ARM I assume it’ll need to be an actual fat binary, with both x86 and ARM code. At that point, unless I actually test the code myself on both architectures, how can I be confident it’s going to work properly?
If you’re using very simple C constructs and not doing anything fancy, it should work, but it’s not clear to me that this approach is preferable to e.g. a Python script. If you’re doing fancy stuff, it’s a bit more chancy as C isn’t memory-safe and has tons of undefined behavior and platform-specified weirdness.
Java is “run anywhere” because they’ve specified the JVM in massive detail and tried to ensure it actually works the same on all platforms. I don’t see how you can have that same confidence if you’re running machine code everywhere.
I guess I just don’t see the use case where this is compelling. If I write a handy Unix utility in C, I’ll just keep the source code around, and compile it as needed.
It might be handy if you need to move such a utility quickly from Unix to Windows, if you don’t have any dev tools set up. But I can’t think of a situation when I’ve needed that.
make -j8 o//tool/viz/deathstar.com
qemu-system-x86_64 -m 16 -nographic -fda o//tool/viz/deathstar.com
You don't have to do anything special either. You just write your TUI program in the conventional UNIX style.I’ll be using this feature in my day-to-day work because I interact with short-lived VMs running the software I’m developing. Being able to run lnav on my devbox and have it live-tailing logs on the VMs running the changes will be an improvement over the current workflow where I manually copy the files onto the devbox.
Can't wait to try this feature.
One use case is where the source and target are different operating systems, and either the target does not have a compiler installed or the source does not have a cross-compiler installed.
I run into this often. I do not wish to install Python or Java on every computer. I have small form factor computers where I need to conserve space. Also, I prefer the size and speed of C over Java and Python.
I think there may be a tad too much "marketing" and trying to be cute in the way this is presented but this work certainly has potential utility, whether or not it is used as intended. Even just reading through the hacks is educational.
Programming does not always have to be commercially-oriented to be useful or interesting.
Agreed!
Thanks for the response, that does sound like an area where this could be useful.
Although (speaking as somebody who doesn’t do a lot of embedded programming these days, so take this with a pinch of salt) I would have thought that if cross-compilation could be made much faster and easier, that might be a better solution overall. I really like what’s being done with Zig around cross-compilation.
Do you mean different OSes? AFAIK, Cosmopolitan isn't cross-architecture.
> I do not wish to install Python or Java on every computer. I have small form factor computers where I need to conserve space.
Now you have to install QEMU on every computer, in which case your options for portability dramatically increase.
Isn't this exactly why you don't see the use case? You're willing to compile.
As someone working on a cross-platform, cross-language packaging tool (https://github.com/spack/spack), it's very appealing to not have to build for every OS and Linux distro. Currently we build binaries per-OS/per-distro. This would eliminate a couple dimensions from our combinatorial builds.
We still care a lot about non-x86_64 architectures, so that's still an issue, but the work here is great for distributors of binaries. It seems like it has the potential to replace more cumbersome techniques like manylinux (https://github.com/pypa/manylinux).
What varies between Linux distros, for your purposes? Different libc? I would naively assume it’s all just ELF so it shouldn’t be that big a deal to make a portable Linux binary.
That seems separate from swizzling ELF/Mach-O/PE, specializing the binary on first run(!), etc, which is all super cool but something I’d be wary of relying on as a solid platform. Maybe I’m being too cautious, though!
Reading the Cosmopolitan docs, it has some really clever optimizations, and I think I’d be more excited about simply a small and fast libc over the flashier APE parts.
Software built with a newer glibc is not compatible with software built with an older glibc. You can't just build an EFL binary on an arbitrary distro.
Moreover, and probably more notably for cosmopolitan, binaries built with glibc aren't compatible with BSD libc or musl. IIUC, cosmopolitan works across the BSDs and musl-based Linux distros, as well.
> Please note that your APE binary will assimilate itself as a conventional resident of your platform after the first run, so it can be fast and efficient for subsequent executions.
I understand there may be no real way out, but this defeats part of the promise/purpose of PAE: assume I use a pae binary, know it's pae, and implicitly share it or copy it to a different machine. But once I have copied it from ~/bin/ it's not pae any more, it's optimized for my OS/arch!
Assuming the first-run optimization step is necessary, would it make sense to provide `--deoptimize` or `--paeize` flag to the same binary, so it's easy to return to the original, without recompiling it from source? Does it lose information when it optimizes? Can that information be tucked away (with an optional flag or env variable) for this step?
What happens if the binary is read-only?
> Cosmopolitan makes C a build-once run-anywhere language, similar to Java, except it doesn't require interpreters or virtual machines be installed beforehand. Cosmo provides the same portability benefits as high-level languages like Go and Rust, but it doesn't invent a new language and you won't need to configure a CI system to build separate binaries for each operating system. What Cosmopolitan focuses on is fixing C by decoupling it from platforms, so it can be pleasant to use for writing small unix programs that are easily distributed to a much broader audience.
- https://news.ycombinator.com/item?id=26271117
... and it only supports x86 (without binary translation), right? It's great to see progress like this, but it's poor form to suggest it's build-once run-anywhere in the same sense that Java is. As far as I can tell, it's not trivial to run these binaries on a RPi.
The tricky part is, those claims are all true, they're just not all true at the same time. The portability claims are true (only if you ignore performance), the performance claims are true (only if you're on x86-64).
The landing page says SIMILAR to Java. It doesn't say runs identically on exactly all versions of every system that Java programs will run on.
Also, this is a C binary we are talking about. Fundamentally, a C binary will not run on multiple architectures (at least not without some sort of special translation mechanism like Rosetta or qemu user-mode emulation).
"Modern desktops and servers" is still pretty much x86-64 for most of the world. Yes, I know there are ARM servers and I'm typing this on an M1 Mac and I own nearly all kinds of Raspberry Pi's (even the minor revs of some models).
The C machine model doesn't dictate native compilation. You could (and in fact IBM OS/400 does[0]) compile C to a binary format that does run on multiple architectures.
Maybe instead of "C binary" you meant "C compiled to native binary", but that linguistic shortcut also encourages a mental shortcut that excludes several interesting design tradeoffs. In the IBM TIMI case, the equivalent of setting the execute bit on the binary triggers OS/400 to perform the native code generation and save that native code to disk, similar to some versions of Android Runtime, but without mandatory garbage collection.
It's extremely impressive, and dare I say useful for several cases, but it isn't really like java unless you squint really hard and pretend virtualization/emulators like QEMU are akin to a JVM, which is the way I understand that claim is supposed to be taken.
I mean, the analogy works, and I respect why the author believes this is more practically useful than Java, but it's like saying (IMO) that Linux binaries are a universal standard because you can just virtualize or emulate a Linux kernel on the cheap, a la Docker for the Desktop or whatever it's called.
There is actually a discussion on that on this page, I'll embed part of the relevant discussion here:
> It'll be nice to know that any normal PC program we write will "just work" on Raspberry Pi and Apple ARM. All we have to do embed an ARM build of the emulator above within our x86 executables, and have them morph and re-exec appropriately, similar to how Cosmopolitan is already doing doing with qemu-x86_64, except that this wouldn't need to be installed beforehand. The tradeoff is that, if we do this, binaries will only be 10x smaller than Go's Hello World, instead of 100x smaller. The other tradeoff is the GCC Runtime Exception forbids code morphing, but I already took care of that for you, by rewriting the GNU runtimes.
Not exactly what you meant, perhaps, but in the same ball-park
Given Apple has undergone 3 CPU architecture migrations by now and employs Chris Latner, I was hoping they'd move to something vaguely like a modernized version of TIMI (maybe based on LLVM bitcode) as the default XCode target for Apple Silicon. Rosetta has been good enough so far, but I can see a future where Apple starts really specializing cores (say ultra-low power cores for watches and glasses or an extreme form of big.LITTLE) to the point where it makes sense for them to have radically new instruction encodings.
In particular, x86's total store ordering memory model causes some memory fences to disappear at the machine code level. The Aarch64 relaxed memory model allows for lower cache synchronization overhead, but code with correct memory fences compiled to x86 loses this information, requiring overly conservative binary translation/higher overhead TSO mode in Aarch64 binary translators. These days, hardware acquire/release/full flavors of memory fences better match the C++ and Java memory models, but some hardware has load/store/full flavors of memory fences. Binary translation across these flavors means changing all fences to full fences, or else some static analysis that's far beyond anything I'm aware existing at this time.
To be honest, I think the dream of portability which TIMI represented is mostly dead in recent IBM i versions. More and more functionality depends on the AIX compatibility environment, PASE, which doesn't run under TIMI, it is full of standard AIX XCOFF binaries containing POWER machine code. (Interspersed with calls to IBM i-specific APIs which allow PASE binaries to access services provided by code running inside and underneath TIMI.) Given the increasing use of PASE as time goes by, porting IBM i environments to something other than POWER (if IBM were ever inclined) has become closer to being as hard as porting AIX – which is to say, as hard as any other operating system. TIMI has evolved from a genuine source of portability (which greatly aided IBM in the CISC-to-RISC transition) into being little more than a historical vestige and form of backward-compatibility.
Let's see if I can get some inspiration of how things can be built from this repo?
`o/$(MODE)/depend`, a makefile-looking file with the full dependency tree. It is compiled using tool/build/mkdeps.c. So we have a Makefile generator (Justine, you mentioned elsewhere in this thread you didn't want to invent a build system? :)) The Makefile generator is very specific to this project: parses C files and creates that tree.
It is damn fast and, so far, beautifully documented (at least the build parts I looked). Hell, even documentation lines in file preambles are 72 characters wide, and justified. Crazy. Do you manually justify those?
You are also vendoring statically-built GCC, and the folder with executables is <10MB. LLVM C/C++ toolchain is hundreds of megs, compressed.
I am certainly taking inspiration of being in tight control of the compiler toolchain, and beautiful documentation. Not sure I will write my Makefile generator, since my project is also not that big.
Thanks. Cosmopolitan is giving me much more to look at than an αcτµαlly pδrταblε εxεcµταblε.
Each and every build tool should be version pinned so that two build runs today and 4 years ago produce the same (hex neutral) executable.
Otherwise debugging issues becomes a nightmare...
Years ago I had a customer who wanted to freeze a toolset and wanted to be sure that any bug fix they requested changed only the lines of code relevant to the bug (they diffed the binaries and traced each change back to our source change to be sure). They paid an enormous premium for this capability.
They were a major phone switch manufacturer (long since absorbed by someone else). Their original design was, IIRC, a Z8000. As those parts were EOLed they shifted to the 68K and wrote a Z8000 emulator for the 68K. They later shifted to the PPC and ported their Z8K emulator to the PPC. We supplied a frozen version of the GCC PPC cross compiler. They had their own frozen version of a Z8K toolchain I think.
Their SLA was something like "less than five minutes of downtime per decade" -- no rebooting, realtime performance, no other interruption -- and they believed their extreme conservatism helped them get there.
Wow.
Well, if they were willing to pay…
Thanks very much for that wonderful piece of history.
Does this mean their hardware may still be running on a Power PC that is emulating a Z8000 that is emulating a 68000?
But in fact they changed from "running on a 68K that is emulating a Z8000" to "running on a Power PC that is emulating a Z8000".
They ported their Z8K emulator from 68K to PPC which wasn't super hard. When the PPC was designed it was planned as an upgrade path from 68K series (remember it was a JV between Apple, Motorola and IBM; the first two, at least, had vested interests in making that transition as easy as possible).
This whole stack of emulation sounds crazy but given their needs, it wasn't.
It's not done usually because it's also very useful to take advantage of new language and library features and all software is supposed to not break compatibility with new versions, so if there is an issue due to upgrading it's because an incompetent maintainer failed at their job.
There's 132,059 lines of Makefile code that's generated to o/$(MODE)/depend e.g.
o//libc/stubs/gcov.o: \
libc/stubs/gcov.S \
libc/macros.internal.h \
libc/macros.internal.inc \
libc/macros-cpp.internal.inc \
ape/relocations.h
That much code can't be written by hand, and if you don't write that, then your build targets won't be invalidated correctly. You'll end up with a non-deterministic unreliable build, which is much worse than generating some unfancy make that causes the make process to bootstrap itself. Goal is to get those hex perfect reproducible binaries with minimal toil.Also, when you write build configs, do you depend on system-provided tools and libraries? Such as some .so file or the python interpreter? The cosmopolitan mono repo doesn't do that. It currently only requires the make, sh, zip, mv, rm, touch and gzip commands. I'd ideally like to make it more hermetic but so far that hasn't been an issue, since the above tools are so stable.
gcc has a family of command-line options starting with -M than can generate dependencies as a byproduct during compilation. The generated files are in Makefile format, and only need to be included in your main Makefile.
https://make.mad-scientist.net/papers/advanced-auto-dependen...
It simply represents different tradeoffs of convenience vs. correctness. With Bazel, you get correctness but you pay a complexity price.
Is my understanding correct that the binary changes itself when first run on the target platform? That sounds like it will trigger a lot of red lights with many automated defensive mechanisms like anti virus.
I'm sure in the future people will want a more Pythonic build syntax, since with GNU Make it's a bit easy to shoot oneself in the foot. The biggest issue contributors have had so far with the build is that we need to specify in the top-level Makefile a correctly topologically-ordered list of `include foo.mk` lines. Otherwise some pretty counterintuitive errors happen. But one grows used to it after a little pain and the repo feels like second nature. Ultimately, I just really didn't want to invent yet another build system. I'm pretty proud of the fact that (at least for now) I've managed to make GNU Make work so well for such a large repo. Google themselves actually used GNU Make for their codebase until around ~2005 so I'm hoping Cosmopolitan has got at least another decade of use in it.
> Ultimately, I just really didn't want to invent yet another build system.
Is there a reason not to use Bazel initially or in the future when you want a Python syntax?
What I'd like is 2 mechanisms:
First, the ability to disable this optimisation altogether.
Second, a means to run this optimisation as a distinct step, without executing the rest of the binary. For example, if the "--optimize" flag is used, do the optimisation and then exit.
> All you need to do is download the redbean.com program below, change the filename to .zip, add your content in a zip editing tool, and then change the extension back to .com.
> That performance is thanks to zip and gzip using the same compression format, which enables kernelspace copies.
Oh my. Having a web server executable which is a zip archive at the same time is a lovely idea. Have there been any other attempts similar to this?
From the latest issue:
Technical Note:The electronic edition of this magazine is valid as both PDF and ZIP. The PDF has been cryptographically signed with a factored private key for the TI 83+ graphing calculator.
* Java is slowing our tools down
* amalgamation sqlite style is showing the benefits of tightly written, standalone good old C.
* By extension: monorepos are introducing complexity, since they depend on JVM for blaze/bazel.
* x86 is pervasive in our industry
People might also want to consider the cost/benefit trade-off for binary vs source compatibility. If you can code in a programming language that works across platforms, can be readily transpiled to one of the supported statically typed languages with a robust, small and fast toolchain, you have any number of packagers who can quickly make binaries for your platform of interest that makes it convenient to install.You get the benefit of better static analysis vs good old C.
> Please note that your APE binary will assimilate itself as a conventional resident of your platform after the first run...
No way they'd agree on a common binary format.
I suppose this level of portability is more a feature if you're shipping to PCs anyway, though. If you're deploying to servers, you know the arch and OS ahead of time and there is no obvious downside I can think of to just targeting it directly.
That's funny.
.. Add Fabrice Bellard's JavaScript engine to third party
.. Add SQLite to third party
It says in the license:
Copyright 2020 Justine Alexandra Roberts Tunney
ISC License (same as MIT or BSD with unnecessary text removed)