Small C Compilers
bootstrapping.miraheze.org
bootstrapping.miraheze.org
For example, just picking a random segment, you don't have to squint very hard to see that this is a number literal parser:
else if (tk >= '0' && tk <= '9') {
if (ival = tk - '0') { while (*p >= '0' && *p <= '9') ival = ival * 10 + *p++ - '0'; }
else if (*p == 'x' || *p == 'X') {
[...goes on to handle the hexadecimal case...]
(aside - I love the conversion from string to decimal by subtracting the string value of '0', as this will work for any text encoding where the decimal digits are monotonic and contiguous - so ASCII and EBDIC at least...)That's how everyone did character arithmetic since forever, though, especially with the letters. Wouldn't be surprised if it's in K&R. And it probably became a subtle source of errors when environments changed, as mixing semantics and implementation tends to do.
The C spec requires this be the case for both the source and execution basic character sets.
debug; // print executed instructions $ gcc -o c4 c4.c 2>/dev/null
$ ./c4 hello.c
hello, world
exit(0) cycle = 9
$ ./c4 -s hello.c
1: #include <stdio.h>
2:
3: int main()
4: {
5: printf("hello, world\n");
ENT 0
IMM -274739184
PSH
PRTF
ADJ 1
6: return 0;
IMM 0
LEV
7: }
LEV
$ ./c4 c4.c hello.c
hello, world
exit(0) cycle = 9
exit(0) cycle = 26015
$ ./c4 c4.c c4.c hello.c
hello, world
exit(0) cycle = 9
exit(0) cycle = 26015
exit(0) cycle = 10060183
It works on Windows too, if you have gcc installed.Also, see "Adding support for 64 bit targets" commit:
https://github.com/rswier/c4/commit/2feb8c0a142b2e513be69442...
And yes, swieros has much more.
But there's a more realistic compiler inside swieros: https://github.com/rswier/swieros/blob/master/root/bin/c.c that emits actual "binaries" for the emulated toy CPU. It's hard to tell what exactly is going on though, as this is not "very readable and simple code".
From the project description: “A tiny hand crafted CPU emulator, C compiler, and Operating System”
To that end it is organizes into "phases" where the lower levels bootstrap the higher ones, with level zero essentially coding directly in x86 opcodes.
I really want to dig into it more, but for the life of me cannot find the github page again. Any ideas?
https://news.ycombinator.com/item?id=21201413
That article mentions the project I was thinking of:
https://github.com/oriansj/stage0
Amazingly, it seems they are even aiming to boostrap some minimal hardware as well! Super cool.
SLOC Directory SLOC-by-Language (Sorted)
36270 top_dir ansic=35504,sh=460,perl=306
28825 win32 ansic=28716,asm=109
9692 tests ansic=8806,asm=858,sh=28
2395 lib ansic=2252,asm=143
158 include ansic=158
140 examples ansic=140
Totals grouped by language (dominant language first):
ansic: 75576 (97.54%)
asm: 1110 (1.43%)
sh: 488 (0.63%)
perl: 306 (0.39%)
As far as I can tell, the top-level directory is the actual compiler itself. 35k lines is pretty small for a real C compiler that can bootstrap GCC (the versions that were still written in pure C).Compare it with a few others:
8cc/9cc - 10K LOC cproc - 7K LOC lcc - 30K LOC
OK. Let me rephrase: TCC doesn't have 35k lines of code, but with five minutes' work it could be turned into a compiler with 35k lines of code capable of bootstrapping GCC on Linux. That should be enough to compare it to the ones you list.
Anyways, what I was trying to say is that "tiny" in "tinycc" has lost its meaning already.
My CPU is very small, (~300 LUT4 on an FPGA using Yosys), but has a very minimal ISA. It's mostly an accumulator machine with a stack pointer.
(I've heard about chaiscript, but it seems way too big)