Crafting Interpreters
craftinginterpreters.com
craftinginterpreters.com
https://journal.stuffwithstuff.com/2020/04/05/crafting-craft...
EDIT: BTW the book is great, can't recommend it enough! :)
For those who haven't read the post: the test infrastructure compiles (and runs?) the interpreter code as it exists at the end of each chapter in the book. So not only is the code as a whole guaranteed to compile, but if you follow along correctly, at the end of each chapter you are guaranteed to get something runnable.
Throughout I've used both the website and the dead-tree version of the book. The website is great with 2 monitors, but you might be too tempted to copy-paste the code.
The one thing the book doesn't mandate is the use of the Lox test suite, but I think it should be incorporated into the book. It's easier to hack on your implementation when there's a test suite to validate that everything still works as it should at Chapter X.
1. https://github.com/thundergolfer/uni/tree/main/books/craftin...
Hard agree. This is one of my favourite technical books ever. As an amateur language design junkie, I've been through a raft of books. Just counted and there are 10 on my bookshelf, including the Dragon book [0].
This is my favourite. It's beautifully written, very approachable but definitely not trivial. A really impressive piece of work.
[0]: https://en.wikipedia.org/wiki/Compilers:_Principles,_Techniq...
I think that after I finish the interpreter from the first half of the book in this way, I'll go read the C bytecode part without writing any code quickly, and then have a second pass while implementing on the side.
I think I could manage doing Python alongside C though. Implementing in a different language will slow down progress a lot, but I think it greatly helps with comprehension and retention.
If you are curious about languages, I can only recommend trying to learn a new one while following this book.
Either Zig or Rust would be great for the bytecode compiler/interpreter in the second half of the book, which is written with the expectation of manual memory management. I personally chose Zig and it was fun.
That being said, I've implemented both in C# (during peak quarantine seasons) and I don't regret doing both of them as they cover slightly different aspects or do the same things differently. Both books are well written, meant for non-CS programmers, and patiently guides you through small chunks of code, so that you end up with an actual working interpreter. As opposed to other books I've seen where there are long blocks of code explained by long paragraphs, and rely on you looking up the "accompanying source code" to figure out how the code pieces together. Thorsten Ball also acknowledges that he also learned from Nystrom's Wren, so there are some things that follow the same approach (the Pratt parser if I remember correctly). WAIG also documents some clever simple optimizations which is nice to know on top of having implemented Lox.
What I appreciate most from WAIG that is absent in CI is how it re-uses its AST from the tree-walking interpreter to make a byte code interpreter. CI on the other hand, writes the bytecode interpreter from scratch (C instead of Java) and doesn't cover how to emit bytecode given an AST (something I had the fun of doing myself for Lox) as it took the single pass compilation approach (with focus on efficiency).
Anyway, heartily recommend both working through Crafting Interpreters and learning the D language.
Cli, lsp, cicd, docs, reading about other langs
And we arent even talking about the most important piece of code
Personal rant follows. Apologies for being vague about the details, I've also changed some of them.
Our internal product has had 2 different DSLs written for it, one a graphical one written early in the 2000s, and another textual one written about 6 years ago. Both of them were partially driven by the same guy, the first one as an implementer and the second one as a manager. The second one was supposed to replace the first one. That first, graphical one, was disliked by many internal users, and was never actually fully completed.
Many internal apps were written in both those DSLs, and so far we have kept full backwards compatibility with both of them, including still occasionally fixing bugs in the interpreter. We have a vague plan to retire the first (graphical) one in a couple of years, once everyone has migrated to the second one.
I was against this second DSL from the beginning, and had a big fight against that manager guy who suggested it. I wanted to use an existing language instead of writing our own, again. However I was promptly shut out from the decision making process. From what I could tell this was because a guy on another team convinced that manager that writing a new language is trivial, using a commercial parser generator we have the license to.
I might mention that the product has many other features and the language is just one of them. It might be the most user-visible one, but there are many layers beneath it.
So that second language was built essentially by a commitee, without my direct involvement, with another team, which proved to be a long and painful process that took many hours. And to my great chagrin I did end up partially implementing the transpiler for this language and implementing some of framework required to bolt it into the existing product.
This was sort of fun for a while but I was fully aware I was deep into NIH land.
The language itself was simple enough, but many features added around it caused a large amount of bugs. Features included an in-app editor with full auto-complete, customized validations and warnings, automatic refactorings, a simple but performant on-line coverage engine, a primitive debugger, and so on. The bugs were mostly in the interface between our code and the code coming from the second team, since implementation of those features was split between us and them. Also a large number of bugs came from the effort to be fully backwards compatible with the older DSL language, which had some weird features due to its graphical nature, in order to allow conversion from the old language to the new.
The initial rollout of the language was met with some success, but it was mainly users who were relieved they don't have to the first (graphical) DSL anymore. Then users of course started asking for more features in this language.
However, just then, most of the team left. For about a year I was the only one left and I was swamped with work. Now we do have a team back, but it's mostly new people, and each of them is already responsible for an existing feature other than this DSL.
As a result, this DSL has very little documentation, except a tutorial course and some bare-bones help. Also, to cut corners, me and the team lead purposefully made some design decisions that impact performance, making our app very slow at times when users write lots of long functions in this DSL.
But now we're stuck. The current team lead is practically the only one who knows about this language and all the nitty gritty design issues in it. And as a team lead they don't have time to spend on it.
My team lead tries to be continue being an IC (individual contributor), and fix one of more severe of the performance problems on her own, but since they also has to function as a team lead, it took a long time, and also led to a series of very difficult bugs which we still deal with to this day.
I made some effort to wrote a post-transpiler phase which essentially bypasses some of the performance problems users have ran into, making her work somewhat redundant, but it's not quite enough.
At the same time, there are requests from users and internal support to integrate "proper" scripting languages (python) or to expose APIs for our app.
So all this soured writing compilers / interpreters, and DSLs in general, for me. I used to love learning new languages and wrote my own complete and fully tested parser for some small languages, including the C/C++ preprocessor. But now I can't bear thinking about doing it.
Once you have customers for this DSL, with demands and bug reports, and you hardly have the manpower to actually dedicate to improving it, it isn't fun anymore. Not to mention I was against this NIH thing from the beginning, and it pains me to keep supporting it.
Currently I'm thinking about integrating C# as a kind of scripting language, since we're already a Microsoft shop, although that has been changing in the last decade. TBH this was suggested years ago by the same guy who was responsible for the 2 DSLs, but I rejected it then due to performance issues. Now with AOT compilation, Roslyn, and VSCode / Monaco and its extensions, this might just be the correct time. I'm not yet sure how exactly this will play out, but we'll see.