Elixir RAM and the Template of Doom
evanmiller.org
evanmiller.org
BEAM is really a marvel of engineering.
I've mentioned this before, but I've head experience engineers show disbelief at what things it can do -- lightweight processes with isolated heaps, with ability to have millions of them, isolated faults (if one crashes in anyway it won't scribble over memory of others), very low latency garbage collection, awesome monitoring and tracing facilities, live code replacement and so on.
It really feels like having superpowers using it so nice to see a whole new ecosystem of languages on top of it.
Scala's Akka library is another place you can find this (although a bit harder to configure) but I don't know many other languages that offer this.
Not always, because that thread might have written to so some shared data structure and left it in an inconsistent state. Restarting that one thread might seem to work but because there is no guarantee, the safest way is to restart the OS process.
It might not even be your code. Maybe your RPC library or some other package did it internally. You might restart and 99% of time it will work, but it will not be a guaranteed thing.
Now, of course the equivalent counterpart to this is not threads but you can do this with OS processes. And I've done it in Python. You can use pipes, file system writes, sockets to create a reliable system were some processes can crash without bringing down the whole service. But it is cumbersome.
But imagine the sever handles a million connections. Certainly doable in Erlang. Each one handled by one lightweight process. Connection does something unusual, maybe uses a new feature which was added last night. It crashes, that's fine. Other 999999 are ok. If this was a C++ program, it might have segfault-ed and caused the other 999999 connections to be dropped.
I've seen systems which had a periodically crashing subsystem that was being restarted behind the scenes for a while most of the system stayed up without noticeable issues. In business terms that means a smaller ops team, it means not having to wake up at 4am to answer pages (if subsystem heals itself, can just fix it in the morning) and so on.
Moreover, live code hot-patching is often touted as a gimmick. But I've used on a large production clusters while it was serving tens of thousands of requests per second without stopping. Sure it was scary and by that time something has already gone wrong (to need that fix), and it was nice not having to to shut down everything just to add an extra trace or log statement.
Erlang is amazing of course. Just wanted to say that some of those patterns are now available in python and easily usable. Not common to all python code, which is where I think erlang wins out. It puts this stuff front and center, and has first class support. Looking at a random erlang code base, you'll probably see it there. Random python code bases... not so much (ok, queue systems are quite common in Django/Flask projects). It's very rare to see python greelets/eventlets in the wild for example, but generators and async stuff are becoming quite commonplace.
Also python has single dispatch now (built in, not in a third party library). Another thing which Erlang does well (pattern matching). Again, not so commonly used except in modern python shops. These combined with quickcheck for python (hypothesis), and gradual typing really have made modern python a much more happy place. Erlang deserves some big respect for spreading good ideas.
But it makes a difference if it is in one language/framework and built-in. It makes tracing/debugging/developing easier.
Like with Python, yeah can use a queuing subsystems and submit jobs. But that is another service to configure and manage. Can use multiprocessing (and I've done that), but can't launch 1M of sub-processes. Which now changes how you develop. It has a green-thread co-routine support via greenlet (eventlet & gevent) but those share memory and if you do any CPU intensive work will block each other and will also share the heap.
So there's all this nice fault tolerance by default, whereas in Python you have to look for it (and you can forget it)
The key benefit is that a single process has a harder time monopolizing the system and taking over the system as a whole. This allows you to write code where the code is known to be suboptimal because it cannot take over the machine. In contrast, many other systems relies on a proactive model where each line of code has to be scrutinized for utmost correctness, or your system as a whole will fall.
As for the faults, there is a subtle but important difference compared to exceptions: you handle any error, not just the ones you figured could happen. And because resources are bound to processes, when the process goes away, so does its resource. Usually you would have to handle this in other languages.
The final point is that a fault is made transient: the system recovers and reestablishes the last known good state/invariant of the system. This allows you to try again which often handles the fault. And if it keeps failing, the fault is escalated to gradually larger and larger parts of the program.
There is nothing magical about the model. It is just that other languages have next to no convenient support for it.
When without a GUI I use the built-in dbg module. Sometimes via the recon_trace wrapper to get the tracing to be rate limited (if on a production cluster).
ttb built-in modules used for tracing on multiple nodes, capture more data and then save to a file and inspect later.
eflame or eep (not the Erlang Proposal, the profiling tool) + kcachegrind to view results.
Usually we see Elixir compared directly to Go and across the board Phoenix looks like it's not even coming close. I realize it's not the end all of comparisons, but it's just about the best thing we have going at the moment for comprehensive cross-language benchmarks.
tldr; these results are not representative of the framework.
Elixir is fantastic. Erlangs concurrency model takes some time to get used to but is something I've never seen before. Things like state is managed in a separate server process or the ease of building fault tolerant systems with supervision trees.
Plus, for me the best part, Elixir is just "fun" to write. I didn't have this much fun writing software since I first started with programming years ago.
If anyone is looking to get started, check out the book "Elixir in Action". Great book that brings you up to speed with Elixir, OTP, process handling and a little bit of inner workings.
https://github.com/melling/ComputerLanguages/blob/master/eli...
I never had any issues here with sublime
As someone who has also used Erlang in the past I can tell you what's nice about Elixir: super modern developer friendly ecosystem. Mix - the pkg manager - is dead simple and does the job well, the docs are amazing with tons of examples and direct links to source code everywhere. The language itself also has some nice syntactic sugar over standard Erlang, and integrating with Erlang was a breeze. Overall the whole system is just smooth and works. I was impressed.
I should have been more specific: what about Phoenix do you like?
The other bit for me is Elixir's macro support. This lets you do really clean route matching in Phoenix. See http://www.phoenixframework.org/docs/routing for some examples.
and not to throw yet another buzzword at this post but Phoenix is the best framework I've found yet for smaller networked services. so enjoyable to build with.
Most out-of-box apps are too simple, and most web framework app skeletons are insanely overdesigned and, frankly, miserable. Certainly there are a number of big frameworks (django, rails) that can give you most things out of the box, but I'm not a fan of heavyweight frameworks in general. I think Phoenix strikes a great balance
For example, if I have an `user/index.html.eex` template for a `UserView` module, I can render it with:
View.render(UserView, "index.html", name: "colbyh")
At compile time, we simply precompiled the index.html.eex template as a `def render("index.html", assigns) ...` function clause.Now imagine I want a `bio.html` template. Instead of making a template file, I can literally do this in my view:
def render("bio.html", assigns) do
"Bio: name: #{name assigns.name}"
end
...
View.render(UserView, "bio.html", name: "colbyh")
=> "Bio: name: colbyh"
And it Just Works exactly because there's no magic going on. Templates are just function calls, whether they are precompiled from a template engine or defined by the user manually.- templates are just functions
- Views render templates and serve as a presentation layer
1. @var's in templates are not module attributes (though they do seem to be convenient!)
2. The form for `use` is different in Phoenix than it is in Elixir proper – instead of passing opts, you pass an atom that selects the implementation of __using__ you want. (This bit of magic makes no particular sense to me – why wouldn't you just use a submodule, say MyApp.Controller, that has the correct and single __using__ attached?)
This isn't really true. If you check out the ExUnit, you can pass something like "use ExUnit.Case, aysnc: true" and it'll setup your module to run tests asynchronously. You can pass any atom/value combination into use.
HTTPoision.Base, and ExAws to name a few other projects also have the convention of passing options into __using__.
To me, the confusion comes from the fact that the Elixir get-started documentation doesn't say anything about this. It is however listed in the Kernel forms that "use Module" actually just translates to Module.__using__([])
I seem to remember earlier documentation using a lot of render calls directly in the controller, and add to that the fact that view functions were available to both controller modules and templates without any explicit inclusions muddied the waters a bit. your explanation makes total sense though and the documentation is now clearer than I remember so chalk it up to growing pains? haha.
and having templates that precompile to render functions is pretty magical, tbh. would personally love it if the view documentation reflected this on some level.
> having templates that precompile to render functions is pretty magical, tbh. would personally love it if the view documentation reflected this on some level.
You mean like in our nicely formatted ExDocs? :) The docs walk you through similar examples as I've posted here, but go more in depth. https://hexdocs.pm/phoenix/Phoenix.View.html
and thanks for keeping up work on the docs!
What problem domains will Elixir add substantial power for? What use cases do you think it will be a game changer for? Would you consider the difference in the future between Elixir and JS in terms of productivity/power the difference in the 90s between Lisp/Python and Java/C?
With Elixir and the Phoenix Channels library[1], soft real time, bidirectional pub-sub communication is incredibly simple. Startlingly easy, in fact. If this describes "the tricky part" of an application, Elixir is a substantially better choice than other languages right now.
Comparing Elixir to JS isn't fair. Even ES6 is still kind of a shitty language with shitty concurrency primitives. The main reason JS enjoys the popularity it has on the server-side is because lots of people are afraid of learning a new language and everyone knows it from the browser, not because JS is a particularly expressive, fast, elegant, or concurrency-ready language.
That's basically meaningless in the presence of the internet. I'd say "low latency".
Just a bit of a funny comment/aside, but the actor-based message-passing concurrency of Erlang/OTP and Elixir is actually pretty similar to the classic object-oriented programming of Smalltalk, etc.
"I might think, though I'm not quite sure if I believe this or not, but Erlang might be the only object oriented language because the 3 tenets of object oriented programming are that it's based on message passing, that you have isolation between objects and have polymorphism."
Which, yes, is another way of saying "blame C++ then Java".
In general, assuming programmers aren't smart enough to learn a different syntax has always been a winning bet in programming language design.
It's highly productive, especially at scale, relatively bug-free and easy to test (being a pure functional language with immutable data), designed for (and makes almost trivially easy) extreme concurrency without blocking (even by GC), minimizes the risk of crashes due to errors (YMMV of course), nudges the programmer in the direction of sound and productive architectural decisions, is fast "enough" for most non-number crunching uses (much faster than the 90s dynamic languages, slower than Java/C#/Go/Swift), and, although the ecosystem is still small, it makes up for it with probably the best standard library there is in the form of OTP.
I'd say its problem domain lies in the area of large (large codebase and/or large userbase) realtime apps. Or, more generally, it offers its best productivity relative to the competition for applications where most computations are not particularly complex and/or are repeated, but they may need to be performed at a very high throughput. An excellent server/networking language.
Seems a waste to not look at those things. (My guess: it is probably totally outweighed by luck, assuming better languages users are as driven as, say, JS or PHP users.)
That's part of what makes iolists nice to work with, in the face of immutable strings in a strict language; you can make subroutine calls that may return a string or an arbitrarily nested list of strings, and you simply write code that unconditionally gathers "Whatever was returned" into a list, instead of sitting there appending all the time. The common "concatenation" operation becomes a list consing instead of a laborious copy.
I have done some benchmarks of using `writev` vs concatenating smaller strings using memcpy before sending to the network stream (to avoid kernel context switches) and difference is pretty small.
Or you could make your own struct of byte pointers and lengths to make an iovec you can fill with content from []byte's, with the caveat that the caller needs to leave those bytes untouched between when it adds them to the vec and the actual writev. To mix byte and string output, you probably need to use unsafe.
I don't see a way to tie writev in with the stdlib text/template libraries, which stick to plain io.Reader and Writer, without forking them.
Curious how writev compares to the usual approach of just pointing text/template at a *bufio.Writer (which also passes the "generate 40GB of the word DOOM and don't crash" test).
Regardless of the specifics, BEAM and the ecosystem around it is one of the more interesting places to crib ideas from, in that it's pretty different from a lot of stuff out there but has shown some value and longevity in production.
What you could do is create an object that implements, say, interface{ Multiwrite([][]byte) (int, error) }, that also implements io.Writer as a fallback in terms of the Multiwrite so this object is generally useful, but that still gets you no performance win for text/template unless it is modified to incorporate this new interface.
However, that does seem like it's at least a possibility; there are, for instance, already optimizations in net/http for serving up files via the Linux sendfile kernel call, and this would be pretty similar. Providing implementations of Multiwriter for a file and a socket might not be a tough sell; complexifying text/template to actually use it in the core library might be, though. Depends on what kind of performance numbers you can show.
Whether it'd be worth futzing to get templating with less copying depends on whether you can measure much practical benefit over bytes.Buffer, I guess. Takes round tuits to see.
Mostly just thought writev was neat; it seems cool that it's theoretically possible to write Go code using it too.