Cwerg: C-like language that can be implemented in 10kLOC
github.com
github.com
(For all the non-German speakers: "Zwerg" means "dwarf" and "c" is pronounced the same as "z".)
https://github.com/robertmuth/Cwerg/tree/master/FrontEnd/Con...
I think Lisp and C (not C++) share an important characteristic that few other languages share: being low-level.
C is low-level due to minimal abstractions from the hardware.
Lisp is low-level due to minimal abstractions from the AST.
People wanting to program very close to hardware feel comfortable with C. People wanting to program very close to the compiler feel comfortable with Lisp.
Lisp traditionally has these features, but they’re separate features:
* There is a rich syntax for literal data structures
* Code is directly exposed as data structures so it can be manipulated with code (macros)
* Code is written as literal data structures so it has no separate syntax
Common Lisp for example:
(berlin hamburg munich cologne) ; a list of symbols
#(berlin hamburg munich cologne) ; a vector of symbols
(1.3 1.3d6 13 #c(1 3) 1/3) ; a list of numbers
"Hello World!" ; a string
#\H ; the character H
#S(CITY :NAME BERLIN ; a structure (aka record) of type CITY
:COUNTRY GERMANY)
#*1010101111 ; a bitvector
#2A((BERLIN GERMANY) ; a 2d array
(PARIS FRANCE))
Other programming languages have their own syntax for literal data (even data structures like unicode characters, hash tables, unicode strings, decimal numbers, ...)
Common Lisp OTOH has an extensible reader, one can add new syntax extensions for data structures, using so-called reader macros. READ is the function to read s-expressions from text streams. It returns data objects.Lisp has traditionally a two stage syntax
1) S-expressions
2) Lisp
S-Expression level has a syntax to describe data: lists, conses, numbers, symbols, strings, arrays, ...
Lisp syntax is defined on top of that, as s-expressions: variables, function calls, macro calls, special forms (using quote, let, if, progn, setq, catch, throw, labels, flet, declare, ...), lambda expressions. Each macro also can implement syntax.
I'm not sure this is true for the C preprocessor, though, where macros can represent partial structure.
c-mera: https://github.com/kiselgra/c-mera
cmacro: https://github.com/eudoxia0/cmacro
(This one one is implemented in Common Lisp for its semantics, but doesn't use a S-exp surface syntax for the code.)
sxc: https://github.com/burtonsamograd/sxc
(incomplete)
MetaC: https://github.com/mcallester/MetaC
(References Lisp in readme; doesn't use it for implementation or notation, but references ideas. Source code seems to be a core of .c files, and the rest self-hosted in its own .mc language. Somehow provides a REPL.)
My own toy language which is intended to be a C replacement for myself is prototyped entirely in s-expressions.
What else would you use to represent a syntax tree?
https://github.com/robertmuth/Cwerg/blob/master/FrontEnd/Tes...
And here is same the same program in the tentative concrete syntax:
https://github.com/robertmuth/Cwerg/blob/master/FrontEnd/Con...
Different syntax for different concepts, easier for the eye, quicker the understanding, troubleshooting.
rather than
(set (aref a i) (+ i j))
you have set(aref(a i) +(i j))
and no difference between i and i(), and only one type of node, with a tag and zero or more kids, instead of separate node types for having kids and having tagsthe rose tree model feels like a slightly closer fit to the needs of abstract syntax trees, and although it isn't simpler than sexps, it isn't more complicated either
another possibility is the ml approach where juxtaposition denotes function application but the functions are curried so they only ever take a single argument, which usually looks exactly the same as sexps but conceptually associates the other way
set (aref a i) (+ i j)
the only difference is that this is equivalent, which i think is worse in this context: set ((aref a) i) ((+ i) j)
(note that ml only uses this approach for expressions to evaluate. for data, such as asts, it uses the rose tree approach)regardless, the semantics are a lot more important than the syntax
For unit tests I've used pretty-printed JSON. Text editors syntax highlight it plus you can leverage an off-the-shelf JSON library rather than writing your own s-expressions serializer (not that serializers are difficult to write or anything, but just as a convenience).
[1] https://github.com/rui314/chibicc
Cwerg aims to be the best c-like language that can be implemented in 10kLOC. Obviously, best is highly subjective but I want to improve on C not just re-implement it.
What's the goal with this? I mean, where are you going with this?
Research language? Scratching an itch? Learning exercise?
All good answers, IMHO.
But ... is there a real gap in the needs of programmers that Cwerg is attempting to address? If there is, can you explain a little more the actual gap being addressed?
https://github.com/robertmuth/Cwerg?tab=readme-ov-file#inten...
If you can live without them, it can be a replacement for LLVM.
In term of features, Cwerg is roughly in the same space as Odin, Zig, Hare, etc. (see https://github.com/robertmuth/awesome-low-level-programming-...) But I am much more willing to sacrifice compatibility for simplicity, e.g. no shared libs, no varargs, no linking with non-Cwerg code etc.
I also feel that compilation speed has not received as much attention in the compiler space as it deserves. Go-lang was one of the first to highlight this recently.
https://github.com/robertmuth/awesome-low-level-programming-...
From your other comments here it seems your emphasising "understandability by one person", which Oberon as you mentioned was designed to be understandable.
It reminds me of Taylor Troesh's wigwams
https://taylor.town/pardon-2023#wigwams
I need to document my JIT compiler's design which is really straightforward.
I've been loosely reading qbe's sourcecode but I need to go through the bibliography to understand the code more. At the moment it's all unfamiliar and not understandable.
Also, I'd call it "C runtime" not "C-like" of you're not going to have C-like syntax
Just my 2c
You are right in that the surface syntax will not be c-like but the features (or lack of them) will be.
e: misremembered, it apparently wasn't on the front page: https://hnrankings.info/39786663/
It was a somewhat facetious comment, TBH.
Instead of having the whole system be understandable by a single person, each major component should be.
In fact the 10kLOC applies to the frontend and each backend separately which I think is fair as most compiler writers use off the shelf backends like QBE, LLVM or even C.
I am no expert in compilers, but how does it compare to, say...
• TinyC – https://bellard.org/tcc/
• Small-C – https://en.wikipedia.org/wiki/Small-C
• Smaller C – https://github.com/alexfru/SmallerC
…?
It's probably at the right size for exploring some language ideas.
Cwerg aims to be the best c-like language that can be implemented in 10kLOC. Obviously, best is highly subjective but I want to improve on C not just re-implement it.