The BEAM Has Spoiled Me
gvaughn.github.io
gvaughn.github.io
Elixir/BEAM are "better" but better than what? Curious the author's prior language experience.
What's he developing with it exactly? Native apps? Web stuff? Or just learning / experimenting?
Didn't really understand how the Elixir / BEAM approach is different from typical threads, OS processes, or async programming models - maybe asking too much from a short post :). But still would be interesting to have some specific examples of how Elixir/BEAM is better.
- With threads, one thread can take the others down. With OS processes, it's hard to keep track of which ones are alive and available to send to. With (most) async programming models (e.g. Javascript, Tcl) you only end up utilising one CPU. With Ada-style async programming models you have to think hard about deadlocks and mutexes. The BEAM model solves all these problems.
Erlang's basic unit to handle concurrency is called a process which is more or less equivalent to a green thread; it has it's own execution state, heap and stack. Communication to other processes is via sending and receiving messages; there's no way to express shared memory in Erlang/Elixir (although under the hood there is some sharing at times). No shared memory makes process management simpler; if a process crashes or ends there's no dangling shared state for the environment to clean up (although, you can of course have stored references to the now dead process throughout your system). No shared memory, and immutable data structures make garbage collection fairly simple; a process can stop to GC itself, but that doesn't block any other processes. Message passing is asynchronous, which lets you chose at the sender if you want to send a message and wait for the response, or send it and check for the response later; processes are not able to be interrupted, except by death, so the control flow is always clear, you can check for messages wherever you want, but you do have to check. I really like the built in support for hot loading which lets you change your system as it's running; a lot of people don't like that, and we can agree to disagree. :)
Anyway, all this combines to a system where the natural way to write a concurrent system is to have one process per session, and write the process to just listen for messages either from the client or from the rest of the system and process them in a straight forward way. And if you've got millions of connections, you can have millions of processes. Each process tends to be small and easy to understand, but sometimes the interaction between the processes can get a bit tricky to understand, because it just kind of arises out of the message passing.
You can read as much as you want about the topic, but my suggestion would be to go and create something with it. You need to feel the difference yourself.
If you got the time, I really suggest to try the BEAM, you will have fun and you will approach software differently (and this time also better) after you internalise the BEAM way.
"Better" is a matter of what types of systems you're trying to build.
In my 25 year career, I've been paid to work in Lotus Notes, Java, C#, Javascript, Ruby, Clojure, and now Elixir. My employer's system is basically an integration platform. We perform background tasks on behalf of our customers that access 3rd party APIs, our database, and then decide to push some data to other systems. We do have a web interface that is mostly the status of this background processing, plus a UI for when the users need a manual action.
There's other great replies touching on your technical questions, but here's my brief angle on it. BEAM processes are as isolated as OS processes, but they are very quick to spawn and need very little memory. A typical developer-caliber laptop computer can probably support on the order of a quarter million processes. Their isolated nature is supported at all levels of the design of the BEAM.
There are deeper points in what the author says. Even if you let a process crash in the BEAM, but you don't leave the system in a consistent, recoverable state, the BEAM is not going to help you.
Once you realize this, and have a way to apply back pressure, you can get a lot of the benefits of the BEAM in any language. (Not all, being a VM carefully tuned and design for this pattern still means something, but still.)
When you realise that it is so much simpler to leave the system in a consistent state and program very very aggressively, following only the happy path without worrying about errors, you can apply the same pattern every where. Of course within reason that perform can be impacted.
My takeaway is that the BEAM (and the community around it) force upon you a very particular style of programming, and most of the benefits comes from this style, not from the BEAM itself. You can apply the same style in any programming language.
Edit: I remember an interview that I had once with a team that was trying to bring on a senior to improve their PHP code and they were really unused to the code style I used to solve their problem. Their main issue was the way wrote their code duplicated most of the (very large) request data unnecessarily many times and forced PHP to add more chunks of RAM several times as it handled the request. My code cut all of that fat, but because they didn't understand that's what their problem was they wrote it off as "this guy doesn't know how to code". They didn't even ask and I was glad to not move forward, ultimately.
We're not really training enough truly full-stack engineers.
We mostly just don’t care, and too many of us have honed or at least mislearned excuses for our behavior that business folks have no way to talk us out of. Most vexing for me is when customers threaten to leave over performance and the senior staff pull up some graphs that say we have done all we can. It’s fairly straightforward to climb the ranks on such a project by proving them wrong. These sorts often leave 2-3x (that I know how to find, who knows what he real number is) on the table, which is often enough to satisfy the customer. It’s often hard work but there’s nothing magic about it.
To use a specific example, look at establishing independent processes that communicate through message passing. In Erlang/Elixir, it's trivial. 2 callback implementations and you've got yourself a GenServer running independently and expecting messages. A quick function to wrap the message call, and boom, you're off to the races. You want supervision? Bam, 2 short lines of code.
If you try to do that in, say, Python, you'll be bogged down by a language structure that is, by design, encouraging of serial code.
The thing is, the default in the beam is that on crash you are in a recoverable state (file descriptors are closed, memory is freed). You have to work hard to not have that be the situation.
Other languages do not do this for you, and make you manage state back to recoverability with potentially confusing unwinding of stack state.
There's a significant gap between "I write Elixir" and "I leverage BEAM". I'm in the former camp _still_. I just can't find the time/opportunity to dive in and use it to my advantage.
Testament to how performant and easy it is to get things up and running in Phoenix I guess. I believe once this secret sauce is spread far and wide, we'll see uptake in Elixir the likes of which the world has never seen the likes of which.
That's because most tasks we as engineers write rarely have the need for distribution or parallelism. It's mostly "get data A, get data from B, mix, save resulting data C". nd when you need to do that in parallel, you through it in Kubernetes, or Apache Beam.
It's even more pronounced in web/microservices situations because the web server takes care of all of that for you.
[0] https://elixir-lang.org/getting-started/pattern-matching.htm...
[1] https://elixirschool.com/en/lessons/basics/pipe-operator/
[2] https://elixir-lang.org/getting-started/meta/macros.html
Quite a few of my students have said great things about Phoenix in Action, from Manning, and it looks to me like a fantastic resource for beginners. Elixir in Action is also top-notch and will show you what's special about the BEAM.
Great book though.
The auth generator is just a new feature. Learning how the pieces work from the book is still worthwhile.
My advice if you want to avoid confusion as you're learning is to just use the same version of Phoenix the book does! (1.4 in this case)
After you go through it, you can follow the official upgrade guides in just a few minutes.
I liked the book Programming Elixir 1.6 By (Pragmatic) Dave Thomas. Solved the exercises. Did not quite find use for macros. And I have not done any production projects.
The pragmatic studio has some courses for the space: elixir, OTP, LiveView, etc. and free tutorials. The quality is good. Simon Thomson's course on futurelearn was also good. https://pragmaticstudio.com/ and https://www.futurelearn.com/
A recurrent problem for me though has been to write an application. I used to write java's petclinic/petstore application when learning a framework, library, etc. Have not been able to do so with some other ecosystems... perhaps laziness, perhaps have not been able to answer why. And languages in isolation have not clicked for me much apart from maybe making me think in different ways.
Why does this Erlang/Elixir/BEAM interest me then? A colleague who talked highly of it... Then there's Joe, his talks, and Alan Kay's recommendation has Joe's thesis on recommended reading list: https://news.ycombinator.com/item?id=20653453
There is also Erlang Masterclass series 1,2,3 starting here: https://www.youtube.com/playlist?list=PLR812eVbehlwEArT3Bv3U...
Especially for a dynamic language which uses tuples/maps for a lot of arguments.
Imo this is actually much better than an orm because it really hits home that the data in memory are stale as soon as you receive them because database transactions are explicit, and not hidden behind getters and setters.
ORMs can get you into trouble because they try really hard to hide the problem of distributed state and sources of truth.
It has this cute macro-based syntax, and that seems to inform even the "functional" API as well. Yet when you I wanted to do something just a bit more advanced like joining on an association whose name I don't know in advance (or, say, want to parameterize the function with), I had to resort to writing a macro myself and then eval-ing it an runtime.
Well, take the numerous articles extolling the virtues of Sequel, or DataMapper, or, related to the current discussion, Ecto itself. Like how expressive this stuff is. And how ActiveRecord creates a too huge API surface on models, and that that leads to bad design, etcetera. Now to mention the callbacks, which are admittedly problematic.
But then I try those aforementioned libs and see sharp corners like terrible argument validations (in case of Sequel) which lead to silently failing code, or actual hard limitations on what I can ask Ecto to do without resorting to meta-programming.
> I also haven't ever found a good replacement for AR migrations
IME, both Sequel and Ecto migrations are fairly serviceable. But neither provides the schema dumping feature that AR has. Which is pretty handy.
The first part is your schema. It defines an Elixir struct that is your domain model of your data.
The second idea/part is the concept of a changeset. It represents a change you'd like to make to the data store, but haven't yet. These are composable, so you can bundle a bunch of changes into a single, atomic unit, which you then send to your repo. If any part fails, then the whole thing fails, no changes are made to the persisted data.
The last thing is migrations, which strictly speaking are optional, but a huge convience. They define how you'd like your database to be structured. The cool thing about them is that they define rollback operations automatically.
Once Ecto "clicks", it really becomes a strong selling point to the language as a whole. It makes working with persisted data super easy.
For a start, it's very different. That will become self evident if you try to learn either Erlang or Elixir, when coming from a language like Python/C++/etc (and though Elixir looks a lot like Ruby, it's still more like Erlang, tbh). So that can be classed as a negative.
Secondly, fault tolerance in Erlang/Elixir/OTP is not quite as simple as commonly billed. Isolated processes and supervision trees can give you a lot of mileage, but the model is not a silver bullet. I found I still needed to create an error boundary to make an application 'Do the right thing' in very specific error conditions (for example, not killing a running process, when making a single, accidental RPC with the wrong arguments).
Apart from that though, I can't think of much else negative to say about Elixir.
The tools for introspection are amazing. I recently debugged a production issue by opening a console connected to a running system, and calling the functions I suspected to be failing to see what they were returning directly. Not sure if any other language has that facility.
Then there's all the stuff built into Elixir, like the test suite, the build tool, and even an official linter / code formatter.
Erlang interoperability is also a brilliant design decision. Anytime I need to use a library for a task, I look through the Elixir community first. But if I can't find something suitable, there's often Erlang code out there with years of Dev time behind it.
This is just a mere fraction of all I want to say about the language. And most it is very positive.
Clojure has it.
twisted/python has had manhole facility since early 2000s, I still sometimes use it to day
The learning curve has been a little steep and there is a paradigm shift but I've encountered situations where it really ended up helping me navigate through code with precision and clarity. I think Phoenix is the best candidate for a Web Framework as of today, if one desires rapid development. The standard library is rich, documentation is solid, and features like async tests etc are an added bonus.
I'm using Phoenix for our side project and it is swift as a rocket.
Since you're using Phoenix you are already reaping the benefits of the BEAM because Phoenix helps you use good patterns with regard to process isolation. In many web apps you won't really need to deal directly with OTP for quite a while (unless you're doing something rather specialized).
Sorry I should have framed it better.
1. If the call is a short, syncronous operation (like stat() (well, mostly), open(), getpid()), it just calls it directly in your green thread. It occupies the entire OS-level thread while it runs.
2. For long, blocking operations (eg. nanosleep(), recv(), maybe read() depending on the nature of the fd), the runtime will translate your blocking call into a callback in the main event loop, informing your green thread when data is available to read, or your timer has expired, etc. Only once the runtime knows your operation will return immediately will it actually do the call. In the meantime this frees the OS thread to run a different green thread.
(I'm genuinely curious, independent of my newbie elixir evangelism)
But in languages with green threads (like Go), you CAN spawn a million threads and performance will be fine, until you make a syscall.
An example of how this can happen: I once wrote a Go tool that walks the filesystem. I spawned a new thread for every directory, thinking that Go only has ~N kernel threads so performance will be fine. I was shocked to see that, in this scenario, it spawns a kernel thread for every green thread!
There are several tricks that BEAM employs to work around these problems:
- BEAM has dedicated threads to I/O, and system calls will happen on those dedicated threads. So, the green thread aka processes wil not be affected
- (almost [1]) every call that happens "outside erlang" (that is it calls some code implemented in C inside the VM such as regexps etc.) is re-entrant. So the VM can and will re-prioritize tasks and put processes to sleep when needed and will re-start work when the process wakes up or some external work is done and the answer is received back
So in theory all those stat() calls will be scheduled and queued on the separate prioritized I/O thread, the processes will be put to sleep until the result comes back, and then they will be awoken in turn. This may cause problems with the host OS though :) [2]
[1] There are definitely places where code is not re-entrant yet, because you see updates in release notes from time to time, but most code is re-entrant because of reduction counting: https://news.ycombinator.com/item?id=14440205 and https://stackoverflow.com/questions/31751766/reductions-in-t... and because schedulers can steal processes from each other: https://hamidreza-s.github.io/erlang/scheduling/real-time/pr...
[2] A slightly unrelated anecdote: At a previous job due to some improper coding the web server would slowly accumulate up to to a few 100s of GBs of data in memory due to some long-running processes. When the processes were done, the GC would kick in and release that chunk of memory back to the OS. The OS had trouble with quickly freeing and reclaiming that memory. BEAM was meanwhile happily chugging along as if nothing has happened :D
One nit: OTP instead of BEAM. The supervised processes and modules as application is due to the OTP framework. What that means is that, it's possible to have it in other programming languages as well. Or not, and managed the application with some OS process manager.
This resonates with me instinctually, but I’d love to learn more. What disadvantages would there be in jettisoning real microservices in favor of Erlang’s “logical microservices”?
Maybe in very large, mature companies this happy land of backwards compatible, fully independent microservices, deployed totally independently, is possible. But in startups tearing forward at full speed with limited crew it's a pipe dream. The better solution is for proper source control management, so problematic changes that are actually independent can be reverted effectively, leaving everything else alone.
The week N release for microservice A won't have a dependency on a new feature in the week N release for microservice B, because the staging environment for microservice A (at release N) calls into the prod environment of microservice B (at release N-1), so the release-blocking tests wouldn't have passed.
It is a common and critical antipattern to have a single staging environment across multiple services, where each service calls the staging environment of each other service. This is a disaster waiting to happen for just the reason you describe.
With the BEAM, it is easy to create fully independent process trees. The architecture enables it. In fact, you can pull in dependencies and add them as "OTP applications" (which is in essence a fully independent process tree, doing its thing.)
The BEAM lends itself very well to structuring parts of your application as independently supervised process trees.
This is similar to microservices, where each microservice represents a fully independent process tree, with which you can only communicate via well-defined interfaces.
As I said, the author's comparison to Microservices is very unfortunate.
I was mainly taking issue with the author's comparison of processes with microservices:
> You can consider each process a logical microservice
But as you said, the process trees can be considered (micro)services, but I wouldn't say that each process represents a microservice. It's just not accurate in my opinion.
I was really trying to emphasize the logical separation vs. deployment separation and not the "micro" nature. Thanks for picking up on that.
"[LFE] is the oldest sibling of the many BEAM languages that have been released since its inception (e.g., Elixir, Joxa, Luerl, Erlog, Haskerl, Clojerl, Hamler, etc.)"
I find this little example really cool:
https://lfe.io/books/tutorial/concurrent/dist.html
I've drank the Lisp kool-aid when I had to learn Clojure, but LFE is so much more powerful for building highly concurrent and distributed systems due to it being based on BEAM and having zero interop overhead with Erlang, while being a feature-complete Lisp.