Elixir GenServer Explained
papercups.io
papercups.io
On a side note, its pretty amazing to see the impact on readability and understanding that native language features have on something like async. I've worked with similar systems to the example in the past in JS/Go/etc, and while you can definitely do the same stuff presented here, it tends to be a far messier callback-hell ime. For lack of better explanation, this Elixir code just "flows" better to my eye, which I know is a completely subjective and non-helpful unsolicited piece of info. Thanks again.
You achieve concurrency by spinning up processes (spawning functions). This isn't that different to firing an async function in Javascript, or starting a goroutine, etc.
And you communicate between the processes by sending async messages.
Those are the primitives. Asynchronously send messages between processes that are synchronously doing stuff. Go is actually not -that- different, except sending messages is synchronous, too, unless you set a buffer on the channel (and then you have to size it). Though I agree, because Erlang/Elixir keep it simpler than that, and are dynamically typed, it's a lot cleaner to write.
On this particular distinction, in fact, it's one of the benefits the actor model has over CSP. In the actor model (well, Erlang's take on it, but that's one that has been adopted elsewhere such as in Akka), I care about "if this fails, who should be affected by it". I.e., by default, it's just this process, and with a supervisor it'll restart. How do I get it restarting into a good state? If it goes down, who else needs to go down (to keep state consistent)? That's mostly all I have to ask myself; I don't have to ask what all the ways it could go down are (slight caveat there; you still want circuit breakers around external resources).
With CSP, I have to ask "if this goes down, who else IS affected by it". Not who should be, and be intentional about linking them, but determine who accidentally is affected by it. This is everyone who might be sending or receiving across a channel (to borrow Go parlance) this process has access to. Which breaks the abstraction; from the perspective of the current process I shouldn't have to know who is sending or receiving across a channel, but the reality is I -do-. And that's no better than the error model I have in other languages; I have to understand every part of my application to know what may or may not be affected by this one part going down. Go gets 'around' this by just making it so any uncaught exception (err, unrecovered panic) crashes the whole application, which is certainly one way to do it.
But yes, it is way better than callback hell... It just could be better still.
That said you shouldn't write genservers if you can help it. Using the frameworks genservers Is the best choice, Task is a better choice for most cases.
I feel like Erlang/Elixir is designed to handle cases like this robustly, but it's not clear to me how this code avoids losing data when a process crashes, potentially up to 5s worth of updates!
Long Answer - What happens in any language where you batch stuff in memory? If the system breaks down, you lose that data. This isn't unique to Erlang/Elixir. There's a tradeoff that persistent storage is slower than memory, but it's persistent. So do you want to be fast, or do you want to be durable?
However, there is some flexibility to address this at a system level. Namely, you wait for confirmation. A client, be that an actual user, another process, etc, wants to know if the write has been persisted. Whereas many other languages make this sort of caching layer completely transparent to the client (i.e., they return success immediately, and so failures mean invisible data loss as the cache is dropped), Erlang/Elixir's model makes it so you HAVE to think about this. I sent a message to the downstream process; do I care about a response? If I don't get a response for any given interval, I don't know what happened to that message. I can still do useful work in the meantime, but I don't know that my message was fully handled (persisted). I can retry or I can report failure or whatever.
This is true in any distributed system, and in Erlang/Elixir it's expressed very evidently in the language constructs (rather than being hidden from you).
You can build local disk caches if you want (in fact, there are included tools to make this super easy for you; ETS, DETS, and Mnesia allow you to shove Erlang terms into memory storage, disk based storage, and a hybrid of the two with some nice DB-like behaviors, respectively), but you need to choose to slow your message ingestion to the speed of local disk writes in that case (as well as handle synchronization of deletes in the event of multiple writers to your downstream). An depending what your upstream is, that doesn't provide a guarantee (i.e., a write to local disk != a persisted write from a user perspective, because the disk could crash before it ever makes it off the local one).
A little over a decade ago I did a project using Erlang. My next project was using Ruby, and I was suddenly _horrified_ by the fact that I had no idea what would happen if the application crashed (and Ruby apps, at least in those days, crashed fairly often). Erlang/Elixir both forces you to think about these things, and gives you tools to address them, where many other languages (or their libraries) simply assume that we will stay on the happy path.
Of course, you can be very productive using languages that just ignore the possibility of very rare failures. Many successful systems work on the basis that sometimes shit just happens and maybe some data does get lost. But, after programming with Erlang or Elixir for a while, that situation starts to feel less acceptable!
I consider the couple years I built systems in Erlang to be fundamental for me. It's affected, for better and worse, my entire approach to system design, at every level. It's meant the stuff I or (now that I'm in management) my teams tend to write is incredibly resilient (compared with the other teams in the department), but also meant that I have a really hard time with any Silicon Valley interview.
If the process dies, the Supervisor invokes the `terminate/2` callback so you're still able to process events.
Here are the relevant lines: https://github.com/plausible/analytics/blob/b724def948d51a0f...
Keep in mind you need to trap exit signal to tell the supervisor to invoke the callback, as done so here: https://github.com/plausible/analytics/blob/b724def948d51a0f...
The erlang docs also mentions this: https://erlang.org/doc/design_principles/gen_server_concepts...
If I have as significantly long transaction, I can restart the transaction from transaction_id + 1 forward and then handle the actual transaction out of band by hand if its something very odd.
I usually insert a logger in these code paths too, but the crash handler is usually after the logger, which makes it easier to reason what has happened if you are dealing with mutable state.
https://elixir-lang.org/getting-started/mix-otp/supervisor-a...
Can someone chime in on this?
Calls by contrast check out a monitor on the counterparty and crash with a timeout so you have guaranteed delivery and acknowledgement, or crash the calling process.
A lot of times people from other plarforms rush to use casts when they "don't need a response" but the actual meanings have more to do with failure domains and rate limiting back pressure; my personal feeling is you should default to call and only use cast when you need failure isolation.
I'll chime in explicitly to your question though - messages don't just disappear without a reason, but the reason may not be visible to you. It's possible a receiving process got a message, crashed partway through processing it (possibly after you even sent other messages, so dropping those, too), restarted, and now is handling the next incoming messages, making it look like a message dropped. It's possible a receiving process is on another node (in distributed Erlang) and the packet dropped. Etc etc.
These all look the same. The way to handle them is always the same; accept you have at most once delivery, or include an 'ack or retry' mechanism and write your system to handle at least once delivery. As that article mentions, Erlang makes you start thinking of your system as a distributed system from the get go. This initially feels incredibly inconvenient, but as you go down the distributed system, fault tolerant, CAP-bound rabbit hole, it becomes incredibly helpful. The system doesn't hide the things it can't guarantee from you; it forces you to feel uneasy about them.
(Whereas, for example, in Java, I don't know that that worker thread has terminated. So some other thread's fiddling with some bit of shared memory to communicate something to that thread does nothing. Which I've had happen in production and led to a single instance out of the fleet have stale data, which caused us no end of grief. Erlang forces you to ask, as you've just done, "what happens if this message doesn't get received because something happened to the receiver"?)
If you use erlang:send/3 with options or erlang:send_nosuspend/2, you can have some cases where messages would be dropped without trying. I think that may be the source of the drop under load you're thinking of?
Without nosuspend, if you send to a node that's connected, dist will queue it to be sent, but there's no guarantee it's received because networking, and it may not even be sent if the other side has gone away and the send queue is already too large; There is a hard to express constraint that if your message doesn't get received in this case, dist will disconnect eventually, but it's important to note that the tick time-out doesn't guarantee timely delivery either; as long as some data is flowing, you can get quite a backlog; I think I've gotten net_adm:ping times above 30 minutes in some cases. Also, if the dist connection is dropped, but both nodes are online, it will likely reconnect shortly, and some messages will have been lost.
If you send to a node that's not connected, and didn't specify no_connect, dist will queue the message while attempting to connect, but if that attempt fails, the messages will be dropped.
It's also possible for code running as a gen_server (or GenServer, I suppose) to check how many messages are queued for it, and run different logic. I've written gen_servers that would drop optional requests if the queue was large. Also, if client timeouts are known, there are ways to approximate the time spent waiting in queue, and drop requests if they are received after the client already timed out; it's a little tricky to do this though.
So for example one possible implementation of state is that you catch your own crash, save some state and then restore it when you restart
You aren't supposed to ever crash, however the whole language it's based around designing a framework to handle what happens if you do crash
So you design a whole framework of supervisors, watching processes, that all have a defined startup and shutdown order. You build structures with rules such as: if my cache module fails, them restart this whole chunk of application over here.
it's very cool
Obviously, different tolerances to that are going to necessitate different designs. You can configure the size of the message buffer when a GenServer starts, and if you have very low tolerance for lost data, you'd want to use a synchronous message (using call instead of cast, which blocks the sender until it returns) and appropriate error handling.
BEAM has a plethora of features for reliable applications, you just have to apply them appropriately.
I also saw that there is the possibility of doing collaborative real time drawing, but still did not try it.
I have seen that the exported SVGs could be simplified (lots of repeated markup), and I am thinking about giving a try to do a PR.
Excalidraw is superb.
Show HN: Eraser — Excalidraw-based visual meeting canvas - https://news.ycombinator.com/item?id=26330288 - March 2021 (8 comments)
One Year of Excalidraw - https://news.ycombinator.com/item?id=25608336 - Jan 2021 (19 comments)
Show HN: Create a Slideshow with Excalidraw - https://news.ycombinator.com/item?id=25081914 - Nov 2020 (1 comment)
Excalidraw whiteboard – easily sketch diagrams with a hand-drawn feel - https://news.ycombinator.com/item?id=23525648 - June 2020 (54 comments)
Building Excalidraw's P2P Collaboration Feature - https://news.ycombinator.com/item?id=22719576 - March 2020 (1 comment)
End-to-end encryption in the browser - https://news.ycombinator.com/item?id=22663435 - March 2020 (107 comments)
End-to-End Encryption in the Browser and how we did it in Excalidraw - https://news.ycombinator.com/item?id=22655009 - March 2020 (4 comments)
Show HN: Excalidraw – Sketch Hand-Drawn Like Diagrams - https://news.ycombinator.com/item?id=22146973 - Jan 2020 (6 comments)
Excalidraw – a whiteboard tool to sketch hand-drawn diagrams (excalidraw.com) - https://news.ycombinator.com/item?id=22104502 - Jan 2020 (3 comments)
Excalidraw – a whiteboard tool that lets you sketch hand-drawn diagrams - https://news.ycombinator.com/item?id=22101381 - Jan 2020 (21 comments)
At first, I was surprised that the Erlang processes resemble the Actor model in such a minimalistic interface. Then I play with them for a while, I realized I wanted to extract something similar to GenServers. Then when I played with GenServers, soon I was bothered by a bunch of Pids, and yearning for Registry. In the end, Supervisor is also pretty natural because things in the universe generally cannot revive themselves.
In a retrospective, OTP resembles a lot of concepts we are familiar with, the Registry just looks like how the web works because we need to resolve the address problem. And vice versa: Docker, Service Discovery, and Kubernetes look like GenServer, Registry, Supervisor respectively just at different scales.
It seems to me to let our systems fulfill some properties, some recurring themes are required. We'll reuse them or rediscover them one way or another.
"Any sufficiently complicated concurrent program in another language contains an ad hoc informally-specified bug-ridden slow implementation of half of Erlang."
Erlang's been around long enough at this point that for pretty much any problem in Erlang's area of expertice you find yourself solving twice, there's a damn good existing solution somewhere in the OTP.
Oh, gosh. This is how I learned... everything? I feel like I just walked into homeroom without my pants on.
Might be a nitpicking here, but GenServers aren't useful for concurrency. They just manage state, and only process one message at a time. If you're using this as a cache, your reads will be bottlenecked by however quickly the GenServer can handle the read requests.
I can dump 1000s of things into its message queue and then do something else. It'll keep working away.
It's like saying threads aren't useful for concurrency because they can only do one thing at a time.
Sure, maybe not by themselves but typically you would run many GenServers concurrently in your application as part of a supervision tree. Libraries like Broadway (and the underlying GenStage) are essentially just leveraging GenServer to make it easier to orchestrate concurrency and state synchronization across multiple processes in your application. But you could build a comparable system on your own just using GenServer and a dynamic supervisor.
I used to interview ex-coworkers. I don't do it as much anymore because it gets old/depressing and the answers tend not to change that much.
If, for the purposes of thinking about turnover, you look at a former colleague as an 'aggrieved party', we are generally somewhere in the neighborhood of mediocre at agreeing that certain things were the 'final straw', and work backward through that to actions to increase retention.
One of the things that interested me about these conversations was that people tend to work chronologically backward through many of their complaints. But what was surprising was that some people would go all the way back to their first weeks. To the first straws. The warning signs they ignored. In a way this is not unlike talking to someone who just broke up with a romantic partner.
And while their ex may realize that his/her problems started back with some early inconsiderate behavior that snowballed, and resolve to 'do better next time', I rarely see companies do this, unless I instigate it.
Point is, these petty slights stack up, and can become a big part of someone's narrative for abandoning you (often in a huff). First impressions are important, and you dismiss them at your own peril.
The only reason modern computing got where it is today is by standing on the shoulders of giants. If we go back and re-litigate every naming decision of the past several decades just to make newcomers feel welcome, all progress will slow as we collectively devolve into continuous bikeshedding.
(I actually googled what "gen" in "GenServer" was supposed to mean before posting, because I couldn't remember but knew it wasn't "generate", and couldn't find an answer, including in the docs for GenServer)
[EDIT] to clarify, I failed to find the answer in any of the Elixir docs for GenServer. Evidently I should have looked at Erlang.
https://hexdocs.pm/elixir/1.0.4/GenServer.html#content
`The advantage of using a generic server process (GenServer)`
current release (1.11.4), first paragraph;
https://hexdocs.pm/elixir/1.11.4/GenServer.html#content
`The advantage of using a generic server process (GenServer)`
There is the corollary:
If we have not seen farther, it is because giants were standing on our toes.
There's a huge degree of cognitive dissonance on these issues. Just because something had a reason in the past doesn't mean we have to keep doing it forever. Also "nobody understands" why people keep re-inventing wheels. It's not all hubris, at least not all the time. It's also declaring open season on those past compromises, especially the ones people bump into immediately. Especially the ones the current maintainers immediately get defensive about. Technology dies for a lot of reasons, but the apologists never seem to grasp that the apologies are a stopgap. They explain the pain, they don't cure it.Believe it or not, I'm pro-Elixir. It's just that as I learn any new technology, I start filling out proverbial bingo card boxes for the things I can predict the next person will complain about. Is that a bit cynical? It could be, but it also serves as my todo list for when someone makes an offhand comment about how you really don't have to do this dumb thing, you could use this library that does it for you. Now my coworkers either don't have to worry about that or I have a suggestion when they do.
Turns out when people are in pain, they appreciate sympathy a lot more than they appreciate being gaslit about how it's just their poor sense of history. Their blatantly implied selfishness.
GenServer could use a new name. I think we forget sometimes that you can rename old things by giving them a second, better name and phasing out the old one. I don't think there's anything seditious in that statement.
Designing a programming language is hard, especially when building on top of a 35 year old language like Erlang. As is often the case: if engineers could change the past without breaking everything in the present, we would have already done it. :)
Anyway, it's just about getting used to the terms and what they truly mean.
Remember that Erlang, too, and some of these technologies, names, and conventions, are pretty old and that they may not have been as evocative then as they might be today with some of their conventions. "Process" almost certainly would have been confusing to the neophyte, but appropriately descriptive nonetheless, but "Gen" for Generic.... eh.
Anyway I guess the point is that I have to look inward to see if something like this rubbing me the wrong way is the substandard choice of the project I'm diving into, or if it's me bringing unwarranted biases and assumptions to the table.
That said, this is one of very few criticisms I can level at Elixir, and it's a tiny part of the eco-system - literally, one single name. I wonder would it be easy to alias?
alias GenServer, as: GenericServer
Done.
Examples: WebAssembly(neither web, nor assembly), Serverless (has a server, actually), JavaScript.
IDK, I don't live in Elixir and that name trips me up every damn time I need to read/write some. Everywhere else, "gen" is typically a shortening of "generate" (which I don't love either—just write the word—but it's fairly common), and "generic" rarely occurs in code at all (elsewhere related to programming, yes—in code, no)
[EDIT] incidentally, as long as I'm complaining about Elixir, I've done a lot of Ruby and have no clue whatsoever why people act like Elixir is similar to it, yet constantly see "oh yeah, it's so easy for Rubyists because it's so similar". Then again, I haven't done any Phoenix with Elixir, so maybe they just mean Phoenix is Rails-like.
> “I'm ashamed to say most of my Elixir education has been through trial and error, figuring things out as I go along“
This attitude is so prevalent in software as if we are all supposed to be divined with programming knowledge the moment our IDE spins up. In every other industry that’s exactly how you learn: Get your hands dirty, make mistakes, and fix them.
There’s a lot of things wrong with software engineering - harboring this attitude that self-taught learning is bad or shameful makes it unnecessarily worse
Great write-up otherwise
This is an interesting question.
There's nothing wrong with being an autodidact, per se, but in my experience in hiring people and working with autodidacts: Autodidacts are not very good at recognizing their own limitations[0] and they usually do not have a lot of breadth of knowledge outside their area of interest. This can even go so far as not knowing that there is such a thing as "Entire Field of Study Foo". (Let's say Foo is Discrete Mathematics.)
Will they be bad hires? It depends on what you need -- and they can absolutely perform as well or better than any "more educated" person. It very much depends on the actual person... as almost all work does.
[0] This explicitly ties into the second point about breadth of knowledge. It's about knowing what you don't know and are willing to research further. The problem with autodidacts is often that there's such an insane leap from "oh, i know this" to a seemingly totally unrelated field which has deep connections to what you're doing. That's what a "broad" education in CS is there to teach you. Are you going to use all of that? Hell no! But just maybe that 5% you do end up using makes up for the 95% you don't.
You start software development with "Hello, world!" and then you do more, like getting input from the user, storing it in a variable, then eventually you're using classes and working with objects, then you're tying together a bunch of classes and working with APIs.
No one just "studied" how to create a gas turbine and wah-la, it was made. The entire process of literally everything was one learning exercise after another. We started by using rocks as tools. Here we are, having built better tools from experience and a lot of trial and error and fooling around with things.
Hell, Lego are the same exact idea.