Cosmopolitan v3.5
github.com
github.com
I enjoy doing this too sometimes and don't find it too difficult, but damn...
I enjoyed the video. However, both the author and Donald Keith strike as two people who are mainly motivated by aesthetics (which is supremely ironic as they almost define themselves as being the opposite of Loren Ipsum). I like people who deliver results.
I think the author mostly likes the appearance of someone who pays attention to detail. He says that in a movie if it's obvious that the coffee cups are empty, he notices and bothers him. 3 minutes later in his video he has a guy using his laptop at the beach where glare would make it impossible to see anything on the screen.
It was fun, once practised it barely slowed me down, and when people noticed it blew their minds.
Then along came GUIs and proportional fonts and made it all invisible. :'(
https://en.wikipedia.org/wiki/Michael_S._Hart#Writing_style
See examples at http://hart.pglaf.org/ .
/s
Those commits were
72-chars per line.
Which is easier to
do than EXACTLY 80
character per linehttps://tbaggery.com/2008/04/19/a-note-about-git-commit-mess...
Now I am wondering if somehow, some ancient COBOL limit ended up in git because every tool picked it up as convention from an older tool.
[0]: https://github.com/bytecodealliance/wasm-micro-runtime/tree/...
I'd like to someday be able to easily distribute my little GUI apps (for scientific applications) in a compile-once kind of way. Asking users to install a JVM runtime has turned out to be unworkable (everyone installs the Java 8 runtime b/c that's what comes up first on Google and then the apps silently crash/don't-work and users are confused)
> everyone installs the Java 8 runtime b/c that's what comes up first on Google and then the apps silently crash/don't-work and users are confused
Is it not possible to add a tiny check from main()? If version is less than required, then show a (dreaded) JMessageBox or print message to STDOUT/STDERR where users can safely download the correct JDK.Most projects are compiling at JDK11 or JDK17 now at least, if not JDK21, which is the latest LTS.
the latest release is JDK22. so if you’re running 8, you won’t be able to run much modern Java at all.
in the meantime, look at Native Image and jlink, as neither of these require a Java installation on the target machine.
jlink is usually not workable because it requires JPMS. not even guava ships a module info yet, so that’s tough.
Native Image works great, it just needs a cosmo backend. It supports most of Java’s native desktop APIs.
It's a cool hack but I somehow feel like it should not work.
Cosmopolitan is a very cool hack indeed, but IMHO it’s not likely to gain widespread usage.
I haven't played with it, but I think the classic Windows POSIX subsystem used the COFF/PE file formats instead of ELF.
The problem is wrapping the relevant APIs - as you can see w/e.g. Wine. For some functionality - like networking - the surface is pretty small, for others its a nightmare.
If that's possible, why not polyglot executables? After all, an executable file format is essentially another type of source code.
Pair that with a polyglot ABI and polyglot system calls... and you get APE binaries :)
ELF is pretty stable across many architectures and platforms but it's proven to be inflexible enough for plenty of applications where people develop their own, or alter it in varying ways.
I'd rather have a proper interface designed for portability rather than a hacked-together POSIX-but-not-really.
That would make a great slogan for the project.
It's probably a very fun project to hack on but I would advise against distributing the binaries and expecting them to work over the long term.
On Windows, backwards ABI and executable compatibility has always been an extremely high priority, so I think the danger of future breakage is low.
Neither of those speak to macOS, but maybe someone who knows more can help clarify.
Standardization efforts and backward compatibility assumes a well-behaved application. If you explicitly depend on really weird hacky stuff like abusing corner cases in the object file format, you risk breakage and you will break.
In the Unix world, firstly, platforms that aren't Linux typically say that if you don't do syscalls through system libc, all bets are off. Second, if they standardized a few things here and there, they are likely standardizing stuff that will formed applications linked with typical libraries will exercise. Standardization does not imply explicitly listing all possible corner cases of your object file format.
On Windows, I happen to be a former Microsoft dev who worked on Windows in 2008-2011. If an app were trying to push the limits of the PE format, I don't think it would get fixed on the platform side... I've seen popular applications do much less bad things and get broken.
No need to be rude.
> Standardization efforts and backward compatibility assumes a well-behaved application. If you explicitly depend on really weird hacky stuff like abusing corner cases in the object file format, you risk breakage and you will break.
Sure, but if you get the standard to include your use then it's not a hacky edge case anymore. Ex. https://justine.lol/cosmo3/ "POSIX even changed their rules about binary in shell scripts specifically to let us do it" > https://austingroupbugs.net/view.php?id=1250
POSIX doesn't sanctify this use.
It was appropriate criticism.
The comment above just said "like, people really value binary compatibility and stuff" -- as if it's OK to code against a very specific moment in time with internal dynamic linker constants and such. No. "It works with how ld.so is written right now" says nothing about standards conformance. Not everything that happens to work is conformant to an ABI. Not every emergent behavior is a feature of an ABI. You would have to not understand what is going on at process load time to think that. Even the highly regarded Microsoft binary compatibility does not work against the model that someone is creating an ELF and PE executable in the same file, they would rightly call that crazy town and mark any bugs as "won't fix".
Is 100% reasonable.
> You simply do not know what you were talking about.
Is where it crosses the line into personal attack. If you just left that part out I think the rest of your comment would be quite good. (I mean, I don't agree with it, but it's a fine argument to make.)
The HN audience at large doesn't know the C world very well and they might mistake your thing for something useful in prod. That's kind of unfortunate.
I also don't have time to dig up that original material and I don't have time to reassess the library to see if it has improved, though I doubt it, because I am still certain it is philosophically an unsound idea. I suspect most experienced C coders if they get one look at that they'd say "ok that's kinda cool but seriously don't do that".
I respect that you know a lot more than I do, and I freely acknowledge I was repeating things I've read from the cosmo author without really understanding the details of how the PE/ELF/Mach-O/etc. formats work.
But my "sense" as an experienced developer is that there really is something here worth pursuing and using -- and that in the worst case, tools built using this will have to reassess their OS compatibility with each new major OS release -- which they kind of have to do already :). I trust the Cosmopolitan maintainers will keep Windows compatibility even if future changes are required. So developers will most likely only have to rebuild with the latest version of cosmocc if the PE loader changes. Maybe Windows developers haven't had to do that thanks to Microsoft's efforts, but it has been a thing on other platforms.
In other words, pragmatically, it would be no additional skin off my nose to have to occasionally rebuild to support future major Windows versions if I get such wide executable portability in return -- for something that would otherwise be supported only on macOS and Linux. Windows developers may feel differently.
But I think anyone who distributes something built in C and whose goal is extreme cross-platform portability/compatibility (and frankly, software longevity due to cosmo libc's future stability) ought to seriously consider APE instead of WebAssembly or creating multiple builds.
Surely that becomes a weaker claim with every year and release that NT/Darwin reshuffle their ABIs and it doesn't break.
For example, on my side I'm developing tooling that allows one to delink programs back into object files. This allows me to commit a whole bunch of heresy according to CS101, such as making a native port of a Linux program to Windows without having access to its source code or taking binary code from a PlayStation game and stuffing it into a Linux MIPS program.
When you're doing that kind of dark magic, it's one of those "know the rules so well you can break them" situation. Instead of following the theory that tells you what you should do, you follow the real-life world that constrains what you can do.
It's rather obvious. It's bandaid that lets us continue on with the current outdated paradigms. Not a real innovation.
I tried rsync for Windows from its "cosmos" pile of binaries (https://github.com/jart/cosmopolitan/releases/download/3.3.1...), and it was too slow to be usable compared to WSL rsync.
For many other utilities I use Busybox for Windows (https://github.com/rmyorston/busybox-w32) which is well maintained, fast, and its maintainer is also very responsive to bug reports.
Would be curious regarding features and performance comparisons with musl, which seems to come up a little short compared to say gnu libc.
Cosmopolitan Libc is 2x faster than Musl Libc for many CLI programs like GNU Make and GCC because it has vectorized string libraries like strlen(). Musl won't merge support it on x86. I even mailed Rich patches for it years ago.
Cosmopolitan's malloc() function is very fast. If you link pthread_create() then it'll create a dlmalloc arena for each core dispatched by sched_getcpu(). If you don't use threads then it'll use a single dlmalloc arena without any locking or rdtscp overhead.
Cosmo has pretty good thread support in general, plus little known synchronization primitives from Mike Burrows. You can do things like build a number crunching program with OpenMP and it'll actually run on MacOS and Windows.
Cosmopolitan plays an important role in helping to enable the fastest AI software on CPUs. For example, Mozilla started an open source project called LLaMAfile a few months ago, which runs LLMs locally. It's based off the famous llama.cpp codebase. The main thing Mozilla did differently was they adopted Cosmopolitan Libc as its C library, which made it easy for us to obtain a 4x performance advantage. https://justine.lol/matmul/ I'm actually giving a talk about it this week in San Francisco.
Wish there was something like musl.cc, i.e., static GCC less than 250 MiB, for cosmo libc.
Also I think there were some numbers showing that it sometimes had better performance than alternatives, but I can't seem to find the post right now.
I use git, I've never felt especially put out that it's a different executable on mac, windows, and linux (situation "a"). I also just use what's common for the OS I'm on since there's so may other differences (C:\foo/bar) vs (/foo/bar) etc... (situation b)
https://github.com/search?q=repo%3Agit%2Fgit+WIN32+OR+mingw+...
If no one had done that work, you'd have to port it yourself.
Cosmo provides a development environment that can ease this work. It's not only about the fat binary.
I don't know to what level C++20 does all this now with filesystem and threads etc... There's also things like libuv, Abseil, etc. But, maybe I'll check out cosmo next.
I like where cosmo is putting effort. The self-hosting toolchain and build system for example. That was always a knowledge barrier for me (I'm a webdeveloper, I know nothing).
Personally, it was so different and shocking, that I started to understand a lot of things by the contrast it created with other code. Cosmos's Makefiles were different, the linker was being used in a way that made me notice it as a distinct part of the toolchain, the demos actually worked, etc. It made me interested in that stuff, made it more acessible.
cosmo make probably works on at least 4x the number of operating systems
# cd /usr/local/bin
# file make
make: /usr/local/bin/make: ELF 64-bit LSB pie executable, x86-64, version 1 (SYSV), static-pie linked, stripped
# stat -c %s make
297104
# cd
# tnftp -4o make https://cosmo.zip/pub/cosmos/bin/make
# stat -c %s make
1397059The opportunity for data corruption just freaks me out.
0 - https://news.ycombinator.com/item?id=40040342
1 - https://docs.pex-tool.org/
2 - https://shiv.readthedocs.io/en/latest/
It's already being used for serious things - Mozilla's llamafile and redbean are two examples that come to mind.
Just go through it, it answers all questions you may come up with IMO.
Or you mean about embedding Lua? The official docs are great too: https://www.lua.org/pil/24.html
If you prefer a tutorial, there's plenty as Lua has been around for a long time. Search on Youtube for example and there's lots of videos.
The top hit on DDG for me was https://lucasklassmann.com/blog/2019-02-02-embedding-lua-in-... which seems pretty straightforward to follow.
[1] https://bellard.org/quickjs/ [2] https://github.com/saghul/txiki.js/
It's amazing, but I should warn documentation right now is not quite thorough
- Relatively obscure language, so potential contributor base is limited from the start
- Lua shares PHP's "One datatype to rule them all" design, which works but feels ugly. In PHP it's "arrays", in Lua it's "tables", but either way you have all the attendant problems and weird edge cases to learn
- Expanding on above point, Lua's APIs for working with tables are, uh, idiosyncratic. Slicing a "list", last I looked, was an unintuitive monstrosity like `new_table = { table.unpack(old_table, start_index, length) }`
I could keep going, but I have other things to do during my brief time in this universe.
To any Lua aficionados out there, my apologies if I've misrepresented it. Corrections to my misunderstandings will be appreciated.
Lua is among the most used languages in existence.
It's probably the most used language by under-21-year-olds, and almost certainly so for under 16. Roblox is absolutely enormous.
It's the usual choice of embedded scripting language, so a great number of programmers who don't use it as a daily driver, nonetheless learn it to modify this or that program. Being a very simple, minimalist language, this is easy to do.
It's true that it's missing some affordances which you'll find in larger (and frequently less efficient) languages. You'll end up doing more iteration and setting of metatables. That bothers some people more than others. These are the tradeoffs one must accept, to get a 70KiB binary which fits in an L1 cache while being multiples faster than e.g. Python or Ruby.
A lot of people like lua from using it on very small projects, or like, spiritual reasons related to its implementation simplicity. But having used it quite a lot professionally, in practice it ends up being a slog to work with.
All that said, lua is closely tied to C, and embedded in a C project is both its intended use and the place where it fits best. So while I don't really like lua and would almost never choose it myself, this is one of the rare exceptions where I think it's a good choice.
A lot of people want redbean to use TS or python instead, and imo either would make it much bigger, more complex, for relatively little benefit. It would definitely benefit from a more full-featured language, but I think a weirder one that is still intended for embedding in C would be best. Something out there like janet or fuck it, tcl. But lua has a lot of allure for a lot of programmers and I think gives the project a feeling of "old school cool" while still being pretty accessible. So is probably the best choice in the end.
Wait... Could we build Solvespace web version and bundle it into that and have a completely portable CAD program that runs locally but in the browser?
For now it seems to be more like a framework+toolchain than a standalone library.
Justine wrote:
> We've made a lot of progress reinventing the C++ STL.
What is the motivation behind this? Try to reduce (compiled) binary size? Naively, I am surprised that the C++ STL from Clang isn't OK to use, including the license. Or is this a clean room impl thing?EDIT
To be clear: I mean no disrespect with this question. This is an amazing project.
EDIT 2
Ok, I just spotted this commit message:
https://github.com/jart/cosmopolitan/commit/c4c812c15445f5c3...
> We now have a C++ red-black tree implementation that implements standard
template library compatible APIs while compiling 10x faster than libcxx.Can someone explain to an average C++ programmer like me how they make it compile 10x faster !?
"If we `#include <string>` for the LLVM libcxx header containing the `std::string` class, then the compiler needs to consider 4,800 `#include` lines. Most of them are edges it'll discount since they've already been included but the sum total of files that get included is 694!"
[1]: https://github.com/jart/cosmopolitan/blob/master/ctl/README....
1. Linux 2.x on little-endian MIPS.
> ... , it reconfigures stock GCC and Clang to output a POSIX-approved polyglot format ...
Did POSIX really _approve_ this? if yes, when?
>The input file can be of any type, but the initial portion of the file intended to be parsed according to the shell grammar [...] shall not contain the NUL character.
But the initial portion of, https://cosmo.zip/pub/cosmos/bin/cat for example, does contain the NUL character.
main jart@luna:~/llamafile$ hexdump -C cat | head
00000000 4d 5a 71 46 70 44 3d 27 0a 0a 00 10 00 f8 00 00 |MZqFpD='........|
00000010 00 00 00 00 00 01 00 08 40 00 00 00 00 00 00 00 |........@.......|
You'll notice there's no NUL characters on the first line, and that the subsequent NULs are escaped by a single-quoted string, which is legal. The rules used to be more restrictive but they relaxed the requirements specifically so I could work on APE. Jilles Tjoelker is one of the heroes who made that possible.You can write to Austin Group mailing list and ask for clarification if you want.
Cosmopolitan is a libc replacement. Your program is still compiled to real CPU instructions, which execute much faster.
I guess not.
There's way too many readme's that just jump in and assume the reader already is familiar with the project. it's a pet peeve of mine
It makes two architecture fat binary, but removes all code not needed. Resulting binary is typically only 30-50% larger than 1-architecture binary.
Go to https://cosmo.zip/pub/cosmos/bin/ and download an executable, like https://cosmo.zip/pub/cosmos/bin/basename .
% curl -O https://cosmo.zip/pub/cosmos/bin/basename
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
100 663k 100 663k 0 0 440k 0 0:00:01 0:00:01 --:--:-- 441k
% chmod +x basename
On macOS with an M1: % arch
arm64
% sha256sum basename
2e4cf8378b679dd0c2cdcc043bb1f21eaf464d3c3b636ed101fbc5e3230eb5fb basename
% file ./basename
./basename: DOS/MBR boot sector; partition 1 : ID=0x7f, active, start-CHS (0x0,0,1), end-CHS (0x3ff,255,63), startsector 0, 4294967295 sectors
% ./basename /dev/null.txt
null.txt
On FreeBSD with an amd64: % uname -sm
FreeBSD amd64
% sha256sum basename
2e4cf8378b679dd0c2cdcc043bb1f21eaf464d3c3b636ed101fbc5e3230eb5fb basename
% file basename
basename: DOS/MBR boot sector; partition 1 : ID=0x7f, active, start-CHS (0x0,0,1), end-CHS (0x3ff,255,63), startsector 0, 4294967295 sectors
% ./basename /dev/null.pdf
null.pdf MZqFpD='
ø @ ' <<'justine0e5c7z'
È ²@ë ëëHƒì1Ò½ ëéE ü‡>à¿ p1ÉŽÁúŽ×‰Ìûè ^îr ¸ PP 1ÿ¹ ó¤‡ÒÿêŽ ŽÙ¹ ¸P ŽÀ1À1ÿóª€ú@tè °1É0ö¿Pèo ŒÆƒÆ ŽÆOuóê ' SR´Ís1ÀÍrF¸¹ ¶ » ŽÃ1ÛÍr3´Ír-ˆÏ€ç?€áÀÐÁÐÁ†Í1öŽÆ¾‡÷¥¥¥¥¥¤“«®‘«’«Xª’[ÃZ€ò€1ÀÍr÷ë¡PQ†ÍÐÉÐÉÁ1Û°´ÍYXrþÀ:$v
°þÆ:6)v0öAÃP1ÀÍXë̉þ¬„Àt » ´ÍëòÃW¿¹&èèÿ_èäÿ¿Á&èÞÿóëü¹ ¾ …Àt
QV—¾Ð&è ^YâîÉú…ÒtRV1ɱʬ^ €îZ¬îBIyúà € ÿÿÿ ÿÿÿÿ Uª
justine0e5c7z
#'"
o=$(command -v "$0")
[ x"$1" != x--assimilate ] && type ape >/dev/null 2>&1 && exec ape "$o" "$@"
t="${TMPDIR:-${HOME:-.}}/.ape-1.10"
[ x"$1" != x--assimilate ] && [ -x "$t" ] && exec "$t" "$o" "$@"
m=$(uname -m 2>/dev/null) || m=x86_64
if [ ! -d /Applications ]; then
if [ x"$1" = x--assimilate ]; then
if [ "$m" = x86_64 ] || [ "$m" = amd64 ]; then
exec 7<> "$o" || exit 121
printf '\177ELF\2\1\1\11\0\0\0\0\0\0\0\0\2\0>\0\1\0\0\0vE@\0\0\0\0\0\510\013\000\000\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\100\0\70\0\006\000\0\0\0\0\0\0' >&7
exec 7<&-
fi
exitThe paper for SIGBOVIK 2017 is both a text file and an executable.
Also discussed 11 months ago:
https://news.ycombinator.com/item?id=36821985 (15 comments)
$ file ./bin/python
./bin/python: DOS/MBR boot sector; partition 1 : ID=0x7f, active, start-CHS
(0x0,0,1), end-CHS (0x3ff,255,63), startsector 0, 4294967295 sectors
$ ./bin/python
Python 3.11.4 (heads/pypack1:6eea485, Jan 24 2024, 10:14:24) [GCC 11.2.0] on
linux
Type "help", "copyright", "credits" or "license" for more information.
>>> import platform
>>> platform.processor()
'arm'There's some interesting background on it in this issue: https://github.com/jart/cosmopolitan/issues/141
$ file hello.com
hello.com: αcτµαlly pδrταblε εxεcµταblεEdit: I tried booting it, the issue is still present, I updated the github issue.