A Brief Guide to OTP in Elixir
serokell.io
serokell.io
What a lot of people seem to want is a batteries included framework that does the OTP for you, but doing at a high level of abstraction is difficult. So more commonly you'll see the OTP managed by libraries for more concrete use cases (web servers, database connection pooling, back pressured job processing, etc). Trying to wrap or rebuild OTP from the primitives available in the BEAM is possible, but definitely not an easy task.
Could we potentially have an OTP-like library that cuts down on the boilerplate significantly? Yes. But how much effort would that take and is it much of a win over OTP as-is? Not sure the trade-off is entirely there yet - OTP really isn't that bad. You're still writing less boilerplate than Java in general.
I think we're much better off writing good library-level abstractions over OTP that can be embedded in a supervision tree or used in more specific use cases. If we were to actually attempt some kind of higher level abstraction in the scope of OTP - maybe we can do so with a very different approach like decorated dataflow graphs and runtime property based testing.
To me it looks like a convenience wrapper around Erlang's primitives. I usually use them directly if I need the flexibility to "await" later. If I don't, I have a small function "run_in_subprocess" (mostly to not run into binary gc problems) that spawns and receives within one function.
The only real advantage I see is that it is separately supervised. When avoiding trap_exit, "spawn_link"ing within a process that is itself supervised has a similar effect, though.
Task does the same thing but for a usage model that wasn't previously well supported. It basically removes the last reason I ever did spawn_link directly in Erlang.
2) Task also implements the $callers process library key. What this means is that if you spawn a process with Task, it knows which process was responsible for calling it (note that this is in general different from "the process that supervises it"). In tests, you can use the $callers value to shard global state in a concurrency-friendly fashion.
Examples:
1. Make a mock in Mox. Spawn using Task. Call your mock from the task, Mox knows that the parent test is and serves the "correct mock".
2. Check out a database sandbox. Spawn using Task. Use the database in your task, Ecto knows what is parent test (and checkout) and serves the correct db view.
3. Make an HTTP request from your test. Stuff the $callers parameter into an HTTP header with term_to_binary and hex encoding (probably user-agent is a good choice). Use a plug to put $callers into your phoenix connection genserver. Spawn a Task that accesses your DB. Ecto knows what is the parent test (and checkout) and serves the correct DB view. So the cool thing is that YOUR REQUEST LEFT THE VM and it still worked! and this is composable too, if you do it right.
2) Every OTP behaviour keeps the ancestor info as well, I could use proc_lib instead of spawning directly.
I make more than half of my income from Elixir. That said, the naive receive loop is much easier to understand than any of the GenServer examples. They pollute the module logic with all that handle_* boilerplate. I believe that neither Erlang nor Elixir got the right abstraction there. Method definition in any OO language is easier to understand. At least Elixir should have taken Agent and made it an invisible part of the language. All those calls to Agent in this example https://elixir-lang.org/getting-started/mix-otp/agent.html are still boilerplate.
At the very top of the Extractor readme:
> I don't maintain this project anymore. In hindsight, I don't think it was a good idea in the first place. I haven't been using ExActor myself for years, and I recommend sticking with regular GenServer instead :-)
This both confirms what you said about the Elixir community shying away from macro magic, and makes this library a very bad idea to use in new projects, imo.
In Elixir you don't even have to define all of this boilerplate, just the functions you actually need. Also, the callbacks are not classical methods, a handle_call doesn't have to reply immediately, for example. Distinguishing info, call and cast is also very relevant.
If you wanted something like
{:ok, agent} = Agent.start_link ...
agent.get()
I think you would need a change in the BEAM to recognise that agent is a "particular kind of PID".OTP is well engineered of course, but the basic notion of the spawn -> receive -> loop cycle is so clean and illuminating that I wish newbies would hold out before learning OTP sometimes.
It's natural to think of case-specific abstractions around the primitives that are more germane to the domain at hand than OTP's.
Yes it is neat and its fun to play with but it isn't usually something you want to use in production code. In production code it is important to have processes linked correctly so that errors propagate to callers. GenServer.call handles this for you, and also clarifies intent (blocking call to another process that must respond or fail).
I'm working on a video series about these concurrency patterns; even recorded the first one but the audio is bad so I'm going to re-record it before I publish it when some "real" audio equipment comes in (hopefully this weekend) - let me know what you think:
I knew most of that stuff already so I can't speak to the pure education value, but it made sense, and I think you hit the key points.
I bet some visuals would help it make sense for total nooberz.
AND! I didn't know that about "v" - thanks! :)
They then introduce OTP and explain all the cases it handles.
I've never read a book or tutorial on OTP that just starts with OTP. In fact (a bit tangetially) whenever I encounter abstractions where I don't already understand the primitives I have a very difficult time. Maybe that's just the way my brain is wired though. I have a terrible time with OO programming for that reason. I much prefer a separation of functions and data because it's much easier for me to reason about what is happening.
Anyway, I agree that programmers should start with the primitives, but I've never seen anyone really teach OTP any other way.
Edit: also there are tools such as Erlang.mk[4] by Loïc Hoguin[5] that will handle a lot of the boilerplate for OTP projects and building releases, though it's good to do it manually at least once while learning.
[1] https://pragprog.com/titles/jaerlang2/programming-erlang-2nd...
[2] https://learnyousomeerlang.com/
[3] https://www.manning.com/books/erlang-and-otp-in-action
What a coincidence! I just tweeted about how you can get it and 12 other various programming books for just $8 now: https://twitter.com/AlchemistCamp/status/1311796404830892032
1. OTP behaviors integrate with the OTP supervisor lifecycle management system (i.e. the OTP framework offers the supervisors standardized hooks to start up and shut down your process, guaranteeing that the errors generated during such steps will be in a format the supervisor can use);
2. Processes that implement OTP behaviors will react to `sys` messages; and so can be debugged, hibernated, code-upgraded, etc. on a framework level, "between" the times the process runs developer-defined code, without the developer having to write such handlers into their module.
This is all accomplished by passing control over the receive loop to a framework — `proc_lib` in this case — which in turn passes any messages it doesn't recognize back to your process. It's an inversion-of-control: proc_lib "is" your process; gen_server or whatever is a delegate of proc_lib; and your module is the delegate for gen_server. It's like an OOP class hierarchy in a GUI system, where the base class handles some events, subclasses handle others, and then your module only has to handle the few it's interested in.
But, if there was a way to do exactly that — to specialize one receive statement, defined in some other function, with your own additional clauses — then we wouldn't need the inversion-of-control framework of OTP! We could just write a regular receive statement that "points to" another receive statement as its "parent" or "fallback" (sort of like a chain of firewall rules, each pointing to the next as "what to check next if this one didn't handle it.") Ideally, the compiled result would be one fused receive statement that has all the clauses from all its parents.
And if you think about it, that's totally possible, even if Erlang itself doesn't offer a fancy chainable receive statement like that... because, in a language like Elixir tht has hygenic macros, you can use macros to generate such "fused receive statements"!
I'm honestly really surprised nobody has yet tried. I realized the possibility of this a couple years back, and have been waiting with baited breath for somebody to attempt it. If it worked out, this "tech" could be used to build a wholly-different-feeling language.
I also am not a fan of the complexity and overhead of gen_server so instead I use metal (http://github.com/lpgauth/metal). It's a simple receive loop with an optional init and terminate callback. It implement sys and can be supervised.
I would also like to see "handle_call" and "handle_cast" (maybe even "handle_info" with a blanket noreply-implementation) optional, but until then I can deal with the additional two lines to exit on cast or call if I don't require those.
Maybe if it were instead a macro/parse-transform, injecting all the common "framework" code into the module implementing the server callbacks, you could then at least get the optimizations that result from all the calls into and out of the "framework" being local rather than remote calls. Private functions could be inlined/fused, etc.
Further, if HiPE was made to work again on Erlang 25, such a "statically-compiled" server module could also "stay native", with no context-switches to interpreted code, in a way that HiPE can't presently manage for gen_servers due to the code in `gen_server` and `proc_lib` themselves not executing natively.
The frustrating thing is that when I made exactly such a tutorial two years ago, I was immediately chastised (on my old hosted commenting system) for showing something so low-level to start with:
It's true but in practice this usually doesn't pan out.
For example with just background jobs alone there's the idea of queues, tracking failures / successes, exponential back-off retries, guaranteeing uniqueness, draining, periodic tasks and everything else you'd likely want in a production ready app.
Typically you'd use Redis, Postgres or something else to help with this. Fortunately https://github.com/sorentwo/oban exists and uses Postgres as a back-end with close to 10,000 lines of Elixir.
We use Exq because I've used Sidekiq for years. I should check out Oban, since it could reduce the complexity in our stack (we wouldn't need redis anymore).
Queues: a coordinator genserver process per node, which uses poolboy to only run X jobs at a time
Tracking failures/sucesses: postgres
Exponential backoff: a function
Uniqueness across the cluster: stdlib global locking (https://erlang.org/doc/man/global.html#trans-2)
draining: remote console
periodic tasks: a simple library that knows how to cron
I built a sharded cron-like scheduler in a few hundred lines of code. It's not doing a billion transactions per second, but it's been running in production for a few years without much drama. It's not complicated because it's just postgres and some standard erlang/elixir stuff.
Erlang has plenty of drawbacks, but this isn't it.
Despite the name, mnesia can provide durable storage, without using anything outside of OTP and your nodes' (hopefully durable) filesystems.
Anyway, I do not think Mnesia is a particularly good way to persist stuff, just use postgres.
The main benefit was the ability to load and store Erlang terms natively, along with the fact that it's just a library. The fact that your compute is colocated with your storage lends itself to a different architecture than when you use, say, a stateless web service in front of a relational database.
(Edit: I realize now that I'm replying to a WA old-timer so I'd be curious to know what your assessment of mnesia is with the benefit of hindsight.)
I'm not too worried about a lack of experts and documentation. If you operate at the limits, you'll just need to become an expert, and the source code will be your documentation; mnesia is all written in Erlang, and not too hard to follow; most people won't hit the limits of ets, I'd think. If you don't operate at the limits, you might not need an expert, and if you keep up with OTP releases and hardware releases, the limits get bigger every year.
Looking back, I think we should have spent more time earlier on using the hooks mnesia provides to make healing clusters take less effort; and maybe some automation on cluster expansion. Of course, WA never shied away from manual work by server engineers. :)
(Hmmm, now I'm trying to figure out you are, send me an email if you will :)
Which is why I'm happy others have paved the way.
A huge amount of web apps can get by with a Postgres backed queue that won't break a sweat handling a few hundred writes per second on a low end cloud hosted VPS.
While leveraging something like Oban you don't even need to be an expert. You can just use the library, throw work into it with a reasonable amount of foresight / best practices and you can start developing features in your app.
And if you don't want to put all those pieces together, like you said there's a rich set of modules people have built, like oban. Long rise and live the elixir monoliths ;) (although admittedly even Discord had to swap out some of their most intense functions with Rust, but how many people are writing systems that need to handle millions of concurrents?)
The big difference there is if you have a Dockerized app, Redis won't be restart on every deploy but the BEAM will be so you will lose your ETS cache or whatever you're storing.
Now, I know cache is meant to be disposable but this model of using ETS instead of Redis reminds me of a Python or Ruby app storing its cache in memory. I mean, sure you can do it but for the longest time (7-8 years?) the community as a whole kind of landed on keeping that stuff out of your app's process.
On an unrelated note, I've always been fascinated with the Erlang VM. The idea that you could hot-load modules and have multiple versions of the same module running at once seems really useful. I wonder why other runtimes haven't adopted these features?
Probably mainly due to shared memory. When you know your data can be modified only in one place, changing data structure is easier. Try changing a struct when some other code is using it. It would require a lock for every object in memory.
So there’s this collection of features between the language and VM that support each other, and are generally hard to reproduce outside that ecosystem because they are so interdependent.
Gotta love an introductory article on "OTP" which even includes a section entitled "What is OTP?" but doesn't expand the acronym.
As a newcomer to Elixir's OTP, I cannot divine that experts hold the opinion that the acronym expansion is misleading when that information is deliberately withheld. The first thing this article did was send me off to search the web.
Folks should use a backronym, like "Opinionated Threading Platform".
I don't think it's useful and I think it's pretty dangerous is the first place. You have code that can run two different things, wcgw.
AFAIK Erlang code does not handle any network traffic, all of that is done in C/C++.
Regardless of where the call is handled, you can't just deploy a new switch in a telecom network, but you still want to upgrade them. This is where you really want hot code loading.
1. You can use this during development (where the tight feedback loop is very nice).
2. You can use this during deployment (after verifying that the process will work as intended, hopefully).
(1) is something you could do every day if it was your daily driver language. (2) would happen less frequently (well, as frequently as you deploy) but permit you to do partial updates without bringing the whole system down.
It's worth remembering that deploying Erlang (and Elixir) can be more like deploying with microservices or applications to an application server. You don't want to bring the whole thing down just to change out one part of the system, so you can update just the parts that have changed and need to be updated.
An OSGi "service" is not an actor but, like OTP with servers+supervisors, is the logical unit with which you build your systems. "bundles" (Java JARs with additional metadata) then act as the equivalent "application" that contain the services.
Outside of it being Java, I have strong opinions (positive and negative) that I've decided not to include.
[1] https://en.wikipedia.org/wiki/OSGi [2] https://en.wikipedia.org/wiki/Java_Classloader
If you do this, you'd better make sure that you hotload the new version of B before you hotload the new version of A. Otherwise some process could end up trying to call B:bar through A:foo before the new B has finished loading.
If you rely on hotloading modules as your primary deploy mechanism, you will run into issues like this all the time. So it's not really clear to me that hotloading is a win for your typical stateless web service compared to just deploying a new binary and doing your standard zero-downtime deploy via forking.
It makes more sense when you realize that certain Erlang apps build up a huge amount of state in-memory. In those cases hotloading allows you to load new versions of the changed modules without having to take down the whole VM.
Nope, erlang mailbox is fifo with filtering. When a process sends messages, they are delivered in the same order they were sent.
> Every time the loop runs, it will check from the bottom of the mailbox (in order they were received) for messages that match what we need and process them.
DB development usually uses Ecto: https://github.com/elixir-ecto/ecto, it has OTP servers for connection pooling and other tasks.
For managing background tasks I use Honeydew: https://github.com/koudelka/honeydew
for the lazy: https://github.com/phoenixframework/phoenix
Of course it's always good to understand genservers.
I can see not writing my own GenServer for webapps, but when it comes to concurrency webapps tend to be dead simple. Elixir is a fine choice for these but I feel like its killer features don't shine that much over the alternatives.
Not everything is good tbh, when you see the pain it is to deploy an Erlang app in 2020 ...
I haven’t gotten around to actually trying it myself, but it advertises that it can compile Elixir projects into a single binary you can copy, paste & run (on Linux & MacOS). It’s not the greatest solution, and the mentioned ~0.5s startup time doesn’t make it great for cli tools when compared to Go/Rust/C.
I’m mostly just happy to see there are people out there trying to make Elixir/Erlang easier to use, as much as I love the language, some of the tooling and deployment methods (releases) make me groan.
mix release
Compiles everything, tar/gz's it, sends it to a private s3 bucket. On your server side, you periodically watch the s3 bucket, and when one of your nodes detects a relup, it downloads from the s3 bucket, kills itself. Systemd then restarts it, kicking into the newest version.
That's it. This is maybe about 50-100 lines of code, one external library (pick your AWS library of choice), and one systemd script. I think there's even a library for managing systemd from within the BEAM now.
If you want to be more sophisticated (a rolling blue-green deploy across an erlang cluster) you could probably do it with transactional locks with the :global module in about an additional 100-ish lines of code, including fully verifying the soundness of newly upgraded nodes using telemetry.
Good article, and very worth sharing, but it could have done without that part. That's a terribly unhelpful analogy. For those who Kubernetes it's wildly misleading, and for those who don't it's meaningless.
It strikes me as saying that an oxygen mask in a hospital room is like an airplane: sure they both provide a facility for providing oxygen to the user, but the airplane is significantly more complex than simply an oxygen delivery system. Probably a terrible analogy, just thinking off the top of my head.
For fun, to expound on your analogy, maybe it's like comparing an airport to a hospital. They both feature a person you check in with (triage/ticket counter) to direct you to an appropriate room/gate depending on your needs, but the core mission of each are rather different.
[1] https://kubernetes.io/docs/concepts/workloads/controllers/re...