Chibicc – A Small C Compiler
github.com
github.com
So the chibicc book is not available yet. I'm busy working on the other project (the mold linker) and don't have time to work on it. That being said, I believe the repo is still very valuable for those who want to learn how easy it is to implement a simple C compiler. chibicc's each commit was carefully written so that you can read one commit at a time. I'd recommend starting from the initial commit and observe how each feature (`if`, `for`, local variables, global variables, etc.) is implemented by following each commit.
Why not release mold linker under a dual license (open source and commercial) and sell them a commercial license?
Note that most of their products are available under LGPL, so users are not forced to open source unless they buy a commercial license.
Don't they? I'm surprised to hear that. I have sold a few proprietary versions of my AGPL codes (which were identical to the original one, but with the license stripped). Not enough to make a living, but I can buy some fancy bikes with the money.
In some cases, they even paid --separately-- for support and a few features of the software that were of particular interest to them. For some reason, many companies are extremely frightened of the AGPL, but a dual AGPL/commercial licensing seems to fit them very well. This is a nice model for free software distribution, but it only suits small projects that do not get external contributors.
I've had orgs not bat at an eye paying a few thousand for a software license if there are justifiable productivity returns and the license is required for us to continue to use the software.
If that license is not required, we would never give you a dime. It's sad, but true. Those few thousand, to many companies, are pennies on the dollar. Just price it graciously so you don't leave the little guys out.
You'll get grumbles but do what you need to survive if you want to do this full time and for a living. God knows you, of all people, have earned it.
What I want to see is a program like this, with very simple assembly/object code, and very fast compilation, and uses DynASM to build the executable in-memory.
Then use a tracing JIT to make it fast.
Having good, tiny, reference C compilers is a prerequisite for that, so I'm always happy to see work in that field.
I suspect combining garbage collection, exceptions, closures, tail call optimisation, parallelism, JIT compilation and coroutines is difficult to do orthogonally.
On eatonphil's discord someone recently shared this link: This is a framework for building high performance language runtimes
https://github.com/eclipse/omr
I am currently implementing a programming language and compiler and interpreter in my multiversion-concurrency-control repository.
https://github.com/samsquire/multiversion-concurrency-contro...
I am doing codegen that is interpreted by my imaginary interpreter. My assembly has primitives for thread safe multithreading.
I am using this one:
We could ask for advice from other project developers that have succeeded on their own journey.
- - - -
Fantastic work, congrats and thank you!
Crenshaw's "Let's Build a Compiler" tutorial series --- from the late 80s/early 90s --- uses exactly the same approach:
Binary size of compiler is 300kb where libc is the only dependency.
My test case shows generated binary takes 3x time during execution compared with gcc -O2. Not surprising as it's not optimizing compiler.
I would say excellent.
This makes the stats a little bit darker.
You can't use it for anything other than this compiler. And adding a memory manager would possibly have added 3x more code and complexity than what's here.
[1] - https://github.com/sponsors/rui314
[2] - https://docs.google.com/document/d/1kiW9qmNlJ9oQZM6r5o4_N54s...
I’ve thought about this – why isn’t this a more common thing to do for short-lived programs such as cli programs? Or is it common – could anyone give some examples of well-used programs that do this?
The reasons to not do it that I could think of is:
1. “It’s just bad practice”
2. You may suddenly find yourself having written some kind of malloc bomb, more easily than you think
(And let me take the opportunity to give kudos, this looks super cool!)
Maybe a practical alternative to this—that would still reap the performance benefits—would be to have an allocator that postpones frees until certain time (or e.g. number of bytes allocated) has been passed since the start of the program, and after that point works like normal alloc/free.
Actually this is not far from how garbage collecting languages work..
Having said that nothing stops you from replacing the global allocation functions, so that deallocation is noop. But your program, and especially libraries should still match up malloc/free, new/delete and allocate/deallocate.
Also your default malloc probably does a bunch of bookkeeping that is eventually only used by free. Once you decide that you won't call free or use a noop free, then there are potentially better candidate implementations for malloc as well.
also memory cost money, no big deal for your machine, but scale this to a fleet of thousands of machines and it start to cost big money
then people wonder why their burnrate is indecently high
If you know your program has limited runtime you can treat standard heap allocation functions like an over-engineered arena allocator. :)
Also not arguing for the practice as a general strategy, but feels like it could have a place!
But the prerequisite isn't whether the program is short-lived, but rather that the memory usage pattern is "alloc often, free at the end".
And perhaps, there's something about programming habits. We hear often enough about C having not enough safeties, and one way to mitigate the issue is having "safe" habits. Kind of like how you activate your blinker when you turn even when there's nobody around; if you don't, you might forget to activate it when it matters.
Kind of like I admit I often don't active the blinker when nobody is around, because it means a little inconvenience and it requires you to get one hand off the steering wheel, which might be the only one holding it.
There are code patterns that allow to code like this in a reusable way. Lookup line allocators. You can allocate resources inside a group and release the group in one fell swoop.
Cue one of my favourite comments[1]:
/*
* We divy out chunks of memory rather than call malloc each time so
* we don't have to worry about leaking memory. It's probably
* not a big deal if all this memory was wasted but if this ever
* goes into a library that would probably not be a good idea.
*
* XXX - this *is* in a library....
*/
(Of course, this consideration should be appropriately downweighted by YAGNI, as threading memory management through prototype or internal utility code can by itself easily force it into very non-prototype territory wrt effort.)[1] https://github.com/the-tcpdump-group/libpcap/blob/2180b6e56a...
> portability is not my goal at this moment. It may or may not work on systems other than Ubuntu 20.04.
... that is not good. On the contrary, you should make it so that portability is not an _issue_. IIUC, this should not depend on much beyond libc, or even just libc, so - why should it be Ubuntu-specific?
Packaging for Linux is thankless tedious work.
2. Assuming chibicc doesn't also implement the C standard library, then I doubt that it's an extremely large amount of work.
3. It's not thankless - you are thanked by your users, who aren't forced to get Ubuntu 20.
I cannot imagine saying this to an open source maintainer
You could offer to package and test for other platforms, if it's important to you.