Hacking the Go compiler to add a new keyword
avi.im
avi.im
It would be nice to have more information out there on the internals of the Go compiler. Perhaps there is.
I found stuff like:
- https://github.com/emluque/golang-internals-resources
- https://www.altoros.com/blog/golang-internals-part-1-main-co...
- https://github.com/teh-cmc/go-internals
But yeah, Eli's articles[1] are pretty good.
[1] https://eli.thegreenplace.net/2019/go-compiler-internals-add...
https://github.com/yeokm1/reflections-on-trusting-trust-go
The Gopher Con Singapore (2018) video is a really great summary (20 mins). He modifies the compiler and inserts the backdoor during the presentation:
https://www.youtube.com/watch?v=T82JttlJf60&list=PLq2Nv-Sh8E...
As an exercise, can you help me figure out how to add a token just from the source and discover these quirks?
On second thought, reading from the source and figuring it out could have been possible if you spent hours. But don't you think it should also have some comments to navigate?
With that tool, I prefer to just dive into the source in most cases rather than read documentation, especially when there is a good chance the documentation is wrong/outdated.
Wow, it is pretty cool! Is there an open source software that is similar to this? It reminds me of https://elixir.bootlin.com/. I really want something like these two. A cross-referencer, but seems like the Chromium one does more (editing, for one). The Chromium one seems to do more git-related stuff as well.
Currently checking out https://github.com/bootlin/elixir.
Edit: Oh Wow, GitHub just improved their code search. Pretty cool. Seems like it can do cross-referencing (even between different repositories!).
Take the example of adding a new token. You have to run go generate to generate token strings. But nowhere in the docs or in the code it is mentioned what exactly is the ‘stringer’ and how to install it.
I am also curious about the daily development cycle by a regular Go contributor. How do they make changes, how do they do quick tests before running the whole test suite etc
By the time you've implemented a compiler that "just works" for the language, you notice that you have really inefficient code for those times when all you need to do is increment a variable by 1, given that the hardware has an opcode for that! In order to have nice, general compilation across all use cases, you've programmed the compiler to implement any case of "add X to variable" via the VM commands for "push X's value onto stack, push variable's value onto stack, call add, pop stack into X".
So, I figured I could add "inc" as a keyword. You have the compiler recognize that keyword and translate "inc <var>" into a VM instruction, and then tell the translator how to turn that VM instruction into something that makes use of the opcode.
(Alternatively, you can have it just recognize when it's doing something of the form "<variable> = <variable> + 1", but that's trickier once you've written the whole VM emission step as a single-pass operation.)
I know, pretty basic stuff from the standpoint of a professional compiler programmer, but pretty neat to be able to make an addition to the language like that!
[1] It uses "Jack", a syntactic-sugar-free Java-like language
https://medium.com/trendyol-tech/contributing-the-go-compile...
He'd have fun reverse engineering it.
I went through something similar in 2005, when during an internship I was working on a program reading and displaying generated Python code in which the code comments provided important metadata about the statements on the same line. I hacked the compiler to make the AST generate a new token for comments (they're otherwise dropped), and this was an incredibly useful experience that gave me knowledge about CPython that I still rely on to this day – 16 years later.
This was the only post I could find on internet which talked about Go compiler internals.
Quite a kludgey optimization for the token hash
People to this day laugh at Rasmus choosing the length of function names in PHP decades ago and go does, well, not the same, but let's just say a variant of it. Oh the irony.
This is a difference in kind, not degree.
PHP has wonky function names because Rasmus didn't know any better.
The Go token hash is a private implementation detail; they can change it and only the compiler internals are affected.
Specifically, the weird stuff the author encountered like:
* Generating the token list by parsing the comments of a source file
* only parsing up to a hard-coded token instead of all of the known tokens (?!)
* using a hacky token hashing mechanism that only looks at the first two characters of the token
have nothing to do with Go-the-language.
I think this used to be more true for all languages? Getting defined by the common implementation is something that feels more true today.
That is, is there a particular quality of the Go language you can think of that necessitates only looking at the first two characters of tokens when hashing? Or that requires it to iterate only to `_Var` when generating the keyword map instead of the full length of the list?
If you want your language to actually have some real-world usage, you need real-world performance numbers. Which tends to lead to compiler codebases an few orders of magnitude larger, and much gnarlier to extend.
https://docs.microsoft.com/en-us/dotnet/csharp/language-refe...
To it's detriment, I guess.
bool char decimal double enum false fixed float in int lock long new null ref short sizeof string true uint ulong