Compiling a Lisp: Reader
bernsteinbear.com
bernsteinbear.com
Racket (if you're not aware of what it is) is just a lisp, but the standard library includes an incredible amount of support for creating readers and expanders for programming languages with very little effort.
You aren't even limited to Racket as the target language for your compilation. It's incredibly easy to formulate whatever output you want (like assembly in this article).
Of course there's still an incredible amount of value and fun in doing everything from scratch like this article's series describes. One of the most intellectually rewarding things I've ever done is working through nand2tetris.
You have a great blog /u/tekknolagi.
It compiles to Javacript and Lua, so if you prefer reading those, you can:
https://github.com/sctb/lumen/blob/master/bin/reader.js
https://github.com/sctb/lumen/blob/master/bin/reader.lua
It turns out that you can greatly simplify the code by e.g. "does an atom start with 0x, 0-9, or dash? if so, try converting it to a number."
That turns out to handle cases like 1e9 too. Whereas "1e" becomes a valid lisp symbol. So you can set it to your own value.
The emacs source code is also worth reading. It's not simple, but it's simple enough to be understandable: https://github.com/emacs-mirror/emacs/blob/7b3e94b6648ed00c6...
I like emacs source code because it handles "every possible real-world case that you can think of."
Here is Guile Scheme (not sure if this is standard or Guile specific):
scheme@(guile-user)> 1e9
$2 = 1.0e9
scheme@(guile-user)> (string->symbol "1e9")
$3 = #{1e9}# ;; This the double quote equivalent for symbols.
scheme@(guile-user)> (define #{1e9}# "foo")
scheme@(guile-user)> #{1e9}#
$4 = "foo"
As you can see, 1e9 is read as a number, but that does not mean it cannot also be a symbol, it just requires us sneaking around the reader or using the quoting construct.http://www.lispworks.com/documentation/HyperSpec/Body/02_caa...
https://github.com/JeffBezanson/femtolisp/blob/master/read.c
CMUCL has a fairly complex, but still comprehensible implementation of a Common Lisp reader:
https://gitlab.common-lisp.net/cmucl/cmucl/-/blob/master/src...
It uses a FSM to recognize numbers and symbols when tokenizing, and it demonstrates an implementation of Common Lisp readtables.
Here's how the Common Lisp Hyperspec specifies the Common Lisp reader (as an algorithm):
http://www.lispworks.com/documentation/HyperSpec/Body/02_b.h...
I'm the author and here to answer questions and take constructive criticism.
If you want to comment but don't have an HN account, please check out the mailing list: https://lists.sr.ht/~max/compiling-lisp
From another lisper, congrats and well done!
This one is great though. I just every few days for the next post!
Progress is currently regular-ish because I've already implemented this stuff before I started writing. Once we get to labels and label calls, though, there very well could be a significant slowdown. I haven't figured out how to do that properly!
I'm glad you like the series and I'm very flattered that you posted it.
Unfortunately the internet has moved towards short form, rehashing simple concepts and clickbait titles - there’s so much junk online, and the signal to noise ratio is high.
But people are gaming for most interaction, views, etc and attention spans are shorter. So this is the world we live in
Note, however, that there is heavy lifting done by the default read-table (the CL reader is table-based and is dynamically modifiable[2]), so you will only be able to read tokens if you implement what is on that page.
However, the most common characters are nearly trivial to implement: #\( #\" #\' as well as the most common dispatch characters for #\#
1: http://www.lispworks.com/documentation/HyperSpec/Body/02_b.h...
2: Someone modified the lisp reader to be able to read in valid C89 code, and implemented a backend to compile that to common-lisp: https://github.com/vsedach/Vacietis