With that out of the way, am I understanding correctly that the way this works on Linux/Unix is that the modifies itself (by overwriting the EXE file header with an ELF header)? This seems to have the consequence of making that specific file no longer portable. If I'm understanding things correctly, it also looks like the QEMU hack for non-x86_64 architectures will only work once per file, since after the first time running the file it will no longer run as a shell script on Unix so the QEMU invocation will be unreachable.
Have you considered adding workarounds for this?
As someone who has never messed around with Libc-level programming, it was surprisingly straightforward (and exciting!) to compile Lua all the way.
Cosmopolitan is incredible, thank you so much.
I was wondering if we can accompany runtimes written in C (like lua) with some scripts? So we have cross-platform, double-click-to-run scripting?
For us poor souls who can't write reliable C code, that would be a great thing!
Why is there a forest of files in Cosmopolitan Libc? I tried looking at it on Github to see how things were done, and there were a lot of .h and .S files, I couldn't actually find the C source, though I'm sure somewhere in there is a .c file.
Doesn't fragmenting the source into thousands of files make compilation far slower than it needs to be?
Also, I wonder if/how functions that aren't called in a program get trimmed away by the linker, and thus don't make the executable larger.
Having lots of objects is a good thing because it helps static linking work better. When the Unix linker loads a symbol from a static archive, it pulls in the whole module that defines that symbol. For example, if you define memcpy() and memset() in the same .c file and then your app only needs one of them, they'll both go towards bloating your binary. Workarounds exist like -ffunction-sections and -Wl,--gc-sections but a C library should make assumptions about the fewest flags feasible.
What's the end result? We're able to build executables that are 12kb in size which run on seven different operating systems. The big codebase is what made tiny binaries possible: https://justine.lol/cosmopolitan/howfat.html
I read your post about actually portable executable format but I wonder if it's something that you found immediately helpful for some project or if you work on it just out of curiosity.
build_directory_link := $(shell readlink $(build_directory))
$(if $(build_directory_link),$(shell mkdir -p $(build_directory_link)))Linux will automatically cache all source files in RAM because of the page table cache. Writes to slow storage devices are eliminated through tmpfs. Amazing really.
Also much lower peak memory usage for lld, which is great if you have multiple concurrent link steps.
For a datapoint it takes me ~2 minutes to build a reasonably fully featured ARM linux kernel with a cold cache on a 5 year old i5. wc counts 1399 CC, 65 AS and 411 LD calls. And of course incremental rebuilds are only a fraction of that.
This allows incremental compilation to work much better than it otherwise would. Your first clean build may be slightly slower, but after that, you only have to recompile the files you change.
I was working on something similar (though much smaller in scope), but had to stop when I realised that `ld-linux.so` has some internal APIs that Glibc uses to setup dlopen()/dlsym(); essentially meaning that it is very hard -if not impossible - to load any shared libraries if one does not link to Glibc.
What I wouldn't give for a liblinux, similar to NTDLL.
https://github.com/jart/cosmopolitan/blob/d7ac16a9ed56ebdc70...
Hey I actually tried to make such a thing.
https://github.com/matheusmoreira/liblinux
It provides access to Linux system calls and process start up code that gets all the arguments, environment and auxiliary values. I have several examples of applications written in 100% freestanding C with zero dependencies other than this library:
https://github.com/matheusmoreira/liblinux/tree/master/examp...
It's a bit too low level but I actually planned to make my own ld-liblinux eventually. It currently works with static and dynamic linking. The ld-linux.so is able to link even though nothing else links against glibc. I didn't test dynamic loading though.
I stopped working on it because I found better solution for system calls on the Linux kernel repository:
https://github.com/torvalds/linux/blob/master/tools/include/...
Do you think liblinux could have a future?
Yeah, I have seen that file. Thought last I checked, it crashes in certain conditions [1]; I didn't really go back to it to fix it.
> I actually planned to make my own ld-liblinux eventually
Can you share this code? If not, could you document how exactly would you have done this? I lost all motivation once I realized that ld-linux and Glibc communicate with private APIs, but I'd love to actually work on it if it was possible to do this.
> Do you think liblinux could have a future?
Absolutely. The problem with Linux currently is that the loader rejects programs at the sign of slightest incompatibility before even trying to execute them, which means that there is nothing that the application code could do to work around said incompatibility. If we were able to get our code running (or, as I like to say it, make %RIP point inside the main function), we could then do whatever hacky-bullshit was required make the software work. All I need is to get our code executing, and from that point on, I'll handle system compatibility.
_start:
xorq %rbp,%rbp /* user code should mark the deepest stack frame
* by setting the frame pointer to zero
*/
movq %rsp,%rdi
call liblinux_start
movq %rax,%rdi
movq $__NR_exit,%rax
syscall
static void *after(void *vector) {
void **pointer = (void **) vector;
while (*pointer++ != 0);
return pointer;
}
int liblinux_start(void *stack_pointer) {
long count;
char *arguments;
char *environment;
struct auxiliary *values;
count = *((long *) stack_pointer);
arguments = ((char **) stack_pointer) + 1;
environment = arguments + count + 1;
values = after(environment);
/* start is the main function */
return start(count, arguments, environment, values);
}
> Can you share this code? If not, could you document how exactly would you have done this?Unfortunately I didn't make it that far. I was planning to study how glibc and musl do it just like how I studied their system call implementations. If I start working on this again I'm gonna look this up. I'll need to learn this linking stuff in order to support the kernel vDSO anyway.
What would it take to create an analog of SDL- some kind of lowest-common-denominator interface for mouse/keyboard io, audio output, and software-rendered graphics?
Something like SDL could be ported to run on top of Cosmopolitan but we'd need to port all its dependencies too. That would be a massive undertaking. If Cosmopolitan ends up being the next big thing and attracts a large community of contributors then something like that is bound to happen. Right now it's just a scrappy ambitious project being built for fun. So GUIs aren't something we're able to do quite yet. Although we've got TUIs working great! You can have the best most portable console / terminal apps in the world, today.
Think about it this way. How many years did it take Microsoft and Linux before they could offer a polished GUI? That should give you a rough estimate of how long it takes to develop these sorts of things from first-principles.
It's like telling someone that SICP is the best way to learn Scheme. Maybe, but lots of people learn differently.
Unfortunately I don't have good suggestions myself, but at least they won't feel bad if K&R isn't their cup of tea.
I tried to refresh my memory from previous threads, and I think the general trend has been to recommend (as seen in sibling comments):
https://modernc.gforge.inria.fr/
https://nostarch.com/Effective_C
I've also seen the older:
https://www.oreilly.com/library/view/21st-century-c/97814493...
Mentioned.
There's also Architecture of Open Source Applications, which can help with starting to read some larger code bases, some of which are in C: http://aosabook.org/en/index.html
And there's a general recommendation to go read the source code of the Redis cache/db.
Finally i came across a mention of this short article on gdb (nb: mention of TUI text ui should probably have been in the top, not a foot note):
https://www.recurse.com/blog/5-learning-c-with-gdb
I feel I'm missing a book that has come up often, but can't think of which one.
I did like zed Shaws learn c the hard way, but I'm afraid it's getting a bit long in the tooth: https://learncodethehardway.org/c/
The first three are just useful for getting the basics. The next two are for starting to learn low level stuff. You don't really understand C until you understand how it relates to lower level code.
I thought 21st Century C was good, i've still kept my copy. I'd happily recommend it.
I like the K&R book too - it feels reeeeeeally old but it's really short.
There's a few others that have helped me in various ways but these are a little older -
Love C by Tim Love (free online, my copy is something i just printed out, it's not that long).
Programming from the Ground Up (x86 assembly) - this is available freely online but i bought the book and that helped me a lot with C even though it's a book with only assembly... (to be fair, it does go through calling conventions).
Finally there's another book i love, Advanced Programming in the Unix Environment by Stevens, i have the 6th edition updated by Rago after Stevens' passing. Fascinating book - but huge.
Besides I was in the same boat as you. I come from the world of JS/Python/Go. I even wrote https://github.com/Himujjal/ekon in pure C recently. The reason I thought C would be good is performance and portability. But I found it to be better to invest time in Rust rather than in C. C is a fantastic language but cross platform dependency management is difficult. Unit Testing is also difficult. There are solutions but not as efficient as Rust's ecosystem.
BTW I wish there could be a Cargo for every language.
Odd...I find that a very strange sentiment.
I thoroughly enjoy C, and have even written a compiler for it, but I don't think it's a well-designed language by any account.
I don't like 21st Century C, it may have changed since then but it has a chapter titled "Object-Oriented Programming in C" which is confusing object with Abstract data type.
struct Object;
void Object_init(struct Object *o);
That's ADT.But the next step, which I think is far more important, is to look at the source code of tools that you actually use in real life. Things like cp, or wc, or head. You've probably used them for years without thinking about it, but they're all written in C. Don't look at the modern GNU versions just yet, since they can be packed full of complex functionality, start with micro implementations like Busybox or Toybox. Then look at some OSes that are known for super clean code like NetBSD or OpenBSD. In those OSes you have the benefit of being able to go back to the version that existed 25+ years ago, so you can read through the diffs to see how they adopted new features and found new ways to address C's biggest flaw/risk - memory exploits.
If you're interested in kernel programming, check the Linux and BSD sources, and lurk on the mailing lists for a while to see how people talk about the code. The review process tends to be a bit more brusque than you might be used to on Github, but it's often detailed and results in code that is of a relatively high quality, or at least a consistent standard. It's a great place to learn.
I missed an opportunity to ask in the previous thread: what would it take to link an app in a different language (say, Rust) with this library? Is it enough to just build an object file, that has libc functions assuming LP64 ABI as unresolved exports?
I think your work is outstanding but it looks a bit childish/bizarre with this name, and thus it might prevent some people to trust its reliability.
https://www.youtube.com/watch?v=Xblh12XgQ4o
The first half is relevant.