Enigma: Erlang VM Implementation in Rust
github.com
github.com
I had a lot of fun working on this project, having implemented enough of the VM to run both Elixir and IEx before I stopped.
Ultimately development stalled since I couldn't get any community interest (I was hoping to give an ElixirConf talk but wasn't accepted either). Was hoping to raise some interest and find some contributors in similar vein to https://github.com/RustPython/RustPython
Nowadays I write a lot less Elixir and a lot more Rust.
I believe you'd get much more interest if there was some ambitious new promise for this new VM, such as 10x sequential performance etc.
https://tylerscript.dev/bringing-the-beam-to-webassembly-wit...
There would however be other limitations, like filesystem APIs etc. not being available in the browser that a lot of frameworks in BEAM languages expect, that would severely limit the usefulness, though I guess that applies to either implementation strategy.
The problems with compiling other languages to Web Assembly are primarily 1) lack of goto and 2) inability to instantiate and jump between different stacks. These limitations are especially problematic for languages like Erlang/BEAM and Go because Web Assembly-based VM implementations require an extra level of indirection in order to implement some of their core language semantics, resulting in quite slow performance compared to even a pure, strictly compliant C implementation (and presuming the WASM VM itself adds no overheard, which is not actually the case).
WASM excluded goto support because it was argued that the relooper algorithm required to translate goto constructs to structured WASM statements was sufficiently capable to cover the vast majority of existing code. And they provided evidence to back up that claim. The flaw in that reasoning is that language implementations and similar niche cases have special needs that application code rarely requires, and in that space constructs like goto are crucial to both simplicity of implementation and performance; the inadequacies of relooper become the norm rather than the exception.
Full AOT compilation with dead code elimination seems like a much better fit for minimising that bundle size.
Add a progressive web-app to your mobile's home screen—get a cached WASM bundle, for near-native performance of the resulting "app." Download an Electron app—get a WASM bundle wrapped in a projector.
Even then, there are still browser use-cases where "full fat" apps are no problem. The "apps" in the Chrome Web Store—Adobe has a version of Photoshop in there, I think—make a lot of sense to be WASM. You'll launch them once and keep them open for hours.
I do think that having alternative implementations is good for experimentation though, similar to how Ruby was improved upon ideas from JRuby and Rubinius, even if most users never used those two directly.
And the fact that a LOT of BEAM was terribly undocumented. I'm impressed you got anywhere, personally. The last time I looked at BEAM to understand some weird behavior (that eventually did turn out to be a problem in Erlang/BEAM), it was completely impenetrable.
Your implementation is probably useful as documentation, if nothing else.
Having Enigma, and other implementations like it, provides a huge value in terms of understanding how it all fits together - understanding the Enigma implementation and then going and trying to make sense of the BEAM would probably be a way better path than trying to dive into the BEAM straight away.
https://www.michielrook.nl/2016/11/strangler-pattern-practic...
Probably the language that is most poised to achieve this is Zig; it would be feasible to start by wrapping the entire BEAM in a zig compilation unit; which at the very least potentially offers an easier path to maintaining the codebase across multiple platforms. Followed by hodgepodge doing bits and pieces in zig, which could be achieved via straightforward transliteration at first.
The very different mindset of the rust PL lends itself to total rewrites, which I don't think will sit well in the BEAM community. On the other hand erlang has tons of strange rewrites happening over its own internal ecosystem all the time (gen_fsm -> gen_statem, pg -> pg2 -> pg), etc.
- https://gitlab.gnome.org/GNOME/librsvg completed a migration to Rust.
- https://github.com/RazrFalcon/rustybuzz and https://github.com/immunant/rexpat are making decent progress.
Fwiw this is also a real problem as I am often hearing complaints about windows ABI support in elixir, anyways. Moreover, having written nifs, I have to say the c header for nifs is basically unintelligible without manually parsing the DEFINES because of windows cross-support needs.
Please let me know if I'm using "poised" incorrectly.
The future implied is the completion of the action one is ready for, not further work of preparation for that action.
Your take may, of course, differ.
X509 library: https://hex.pm/packages/x509
Bram verburg talk: https://youtu.be/0jzcPnsE4nQ
Google's gRPC/Stubby has the same problem. It's one reason why gRPC servers are never deployed on the public Internet and always use hefty reverse proxies.
Curious: Did implementing this in Rust expose any bad or interesting behavior when replicating the Erlang language spec (https://github.com/erlang/spec) or whatever reference implementation you were targeting?
It was kind of interesting exploring the OTP internals, especially some of the parts that haven't changed in a long time. One example is the PAM: I think it stood for "patrick's abstract machine" and it would compile erlang terms into bytecode for pattern matches (intended for fast ETS lookups). It's all there in one file and it took a fair bit of digging to figure out how it works since it's been static for a long while and nothing on the internet really documented it.
If this is still the case you should definitely consider contributing to the documentation of those files. Odds are they'll be used by the next person to try something similar. :)
Edit: this book is available for free here: https://www.microsoft.com/en-us/research/publication/the-imp...
I have heard similar before at an Erlang meetup. Are there elements of the VM you encountered that were also static and lacking sufficient documentation? I'm guessing much of this is just tribal knowledge deep inside Ericsson then? It would be great if there were a public repository for these things.
From my own experience, those parts that are commented or documented tend to clarify some specific design constraints (for example, why processes have multiple locks on different parts, and why they are locked in a specific order, or the rationale of the carrier design); but you never really get a clear picture of why things overall are architected the way they are overall, what designs were considered and discarded due to some deficiency, what tradeoffs were made, etc. I think much of the actual content like that which may exist, is either buried in the minds of the original engineers, or in some internal documentation at Ericsson that has never been released. My suspicion is that you'd need to dig through mountains of emails and such to piece together a more complete picture of how things where put together over time.
It's also the fact that the BEAM just has a lot of really complex pieces built in to it after all this time. Everything from binary pattern matching and construction, to garbage collection and memory management, ETS, Mnesia, etc. Each one of those things is not only non-trivial, but have evolved significantly over time, through the hands of many engineers. It also doesn't help that large portions of the C implementation are written in an extremely macro heavy style, which makes it quite hard to read without knowing what all the macros do and how they play together.
Projects like Enigma, or Lumen, have a lot to give back to the community in the form of documenting how these pieces are built. Unfortunately, the lack of a specification for the Erlang language and its runtime, means it is very much a grind to work out how things are currently implemented, and why.
I'm also not saying that the C code is unmaintainable. It's definitely a bear to dive into, but by spending enough time with it, it starts to unfold in front of you. The main issue I have, is that none of the specification/design documentation exists as part of the source repository. Maybe it doesn't exist at all, but in that case, I'd really hope that some of those core engineers would have taken the time to write some of that stuff down. In any case, none of it is readily available AFAIK.
That's what I was originally looking for when I found this.
There is also slice matching on stable, which let's you match on parts of slices: https://github.com/rust-lang/rust/pull/67712/ . It went out in 1.42. It has some stuff which makes binary stuff easier, but not by much. But perhaps someday you'll get native binary matching in the language that's closer to what Erlang offers.
It made it in the 1.42 release.
What they don't realize is that they're often building solutions that are looking for problems, rather than solutions to solve problems. It's also vaguely cultish in the approach.
It's a terrific language and there's a lot of learn from it, but I'd like to see it solve real world problems on its own versus try and screw itself into everyone else's.
In general, though, I think people ought to consider that if they are putting the language they wrote their project in in the marketing blurb for it (given a more serious project), maybe that indicates that the project itself is of little value to other people. "* Written in Rust" isn't a value proposition, it's just an implementation detail. Make real claims about zero crashes, zero leaks, something actually concrete and it can be scrutinized for real.
If you are evaluating a tool/lib/etc that moves at a fast pace and your whole shop is extremely fluent in language X there's huge value add to being able to dive in without a context switch to understand how it works, especially when debugging harder problems.
I don't think it applies to _this_ case where we're getting a VM that is _extremely_ battle-tested. Am I going to use a new OS instead of linux in production because someone tried to write an OS in zig? No. Will I congratulate the author for writing an OS in zig? Yes.
If I am looking for a key-value store and two are equivalent in their purported features and stability, I will choose the one that is written in X that my shop is most fluent in.
After I achieved both of those I decided I'll keep going if there's any interest in the project. There wasn't, so I moved on.
Before even checking the GitHub repo, how are your docs? Can an outsider make a deep dive into your code without needing to pester you with questions?
The author had a threshold of good feedback they needed from the community in a certain amount of time. They got the feedback they needed - people aren't interested in it, probably because of the latter part of your comment.
I don't think that's a valid reason to ask why someone started a thing, people start things for a variety of reasons. As far as I am concerned, they saw the development of a reimplementation of solid tech through and learned a lot from it.
> they already have a VM for BEAM and it works well. Without an additional selling point 'now in Rust' doesn't cut it.
This is spot on though.
I think the goal of Enigma in making the BEAM architecture easier to understand, and providing a great learning platform for getting involved in working on the BEAM itself, or just on Enigma as an alternative is a great idea, and something I know I wish I had been able to have on hand when I was trying to understand the deep inner workings of the BEAM implementation. I if Enigma was only ever that, it would still be worth the effort spent on it.
Lumen does have parts that could likely be shared with projects like Enigma - namely the high-level IR we use in the frontend (EIR, Erlang Intermediate Representation). That IR could be used in any Rust-based VM/compiler targeting Erlang, or an Erlang-derived language, and work there would directly benefit downstream consumers of the IR in terms of better optimization, etc. We've also put work into our term representation, and various parts of the runtime, like reproducing the core parts of the BEAM garbage collector, etc. While some of those things are intertwined with the compiler, much of it could be easily extracted and used elsewhere.
I do think that Lumen is potentially more interesting in the real world usecases though, and I wish you best of luck. Having an alternative Erlang compiler is a lot more complex but allows us to explore new techniques easier that could potentially be ported back to OTP if nothing else.
If nothing else I'd really like to see projects like ours demonstrate the value to the core Erlang/OTP team in addressing the lack of documentation in some areas - ideally in the form of one or more specifications. Erlang deserves a specification at this point - it is very stable, and a spec would at the very least provide additional structure for future evolution. Core Erlang had a specification, but it is very much out of date at this point - considering how widely it is used as an IR for BEAM languages in general, it's disappointing it hasn't been kept up to date.
It was a big help initially to grok the bytecode format.
I thought the very point of Erlang was the distributed nature of BEAM? Failure is normal etc?
BUT
Different versions of code in memory - that sounds like nightmare to debug. Erlang already stores two versions of newly compile code, old one for processes currently actively using that code and new one for all others. Once all processes jump to new code (by exiting from old function module or calling into new version with module:function call) old code is purged.
You can model that today if you script compilation/loading. When loading a module, you could first load it as {?MODULE, ?VERSION} or ?MODULE_?VERSION if we can't stuff a tuple there, and then also load it as ?MODULE, to use when you don't specify a version.
The hard part is deciding what version to call when, and passing that through to the call sites. And also, to figure out how to signal a process that you want it to update its version.
> where you could re-order the message pattern matching at runtime,
Pattern matching order is part of your code, and hot loading is the way to make changes to your code.
> where you can specify arguments to functions in terms of a map, specify the args and the types of that map specification, and have it compile into numbered argument, that way you don't have to add update many many functions to add another argument
You could do this yourself today as well; a function could check if its argument is a Map and demapify the arguments, or you could make a utility call_function(Module, Function, Arity, Map) that demapifies and calls erlang:apply on the function. Or, you could have your rapidly changing functions all just take Maps; I did that in the past with Proplists.
Shouldn't there be a way for the compiler to know enough to convert the map to positional arguments, so long as the map params could be put into a spec? Something like that would be super nice, because I think tail calls in Erlang cost nothing so long as you keep the same arg positions. Having a map is always a convenient way of starting a function and evolving it, but having it with a few more constraints and with equal performance would be better. If I'm not mistaken I remember dart made a similar optimization.
If they're in the same compilation unit, the newer SSA based optimizer might do some cool stuff for you; it's always interesting to look at the optimized code that comes out of compiling/loading.
In general though, I would tend not to worry about efficiency of passing a map instead of traditional arguments. Most likely there's other, bigger, things to worry about; usually finding the right communications patterns and algorithms makes a bigger difference.
(Not necessarily related to Erlang)