The Elixir/Phoenix app rewrite was a little more than 10x faster, much more consistent in response times, and needed far fewer resources to run.
The Elixir/Phoenix app rewrite was a little more than 10x faster, much more consistent in response times, and needed far fewer resources to run.
Elixir and BEAM are clearly enormously better at some tasks but general performance is actually very similar between Elixir, Ruby, and Python if you use similarly performant frameworks.
I think, when it comes to web apps specifically, there are a few advantages that make a practical difference. These seem to be the major ones:
- Compiled templates and more efficient string interpolation
- Improved parallelism and lighter weight processes. This leads to better handling of many requests in parallel (esp. while most are just waiting on IO), better db connection pooling, less blocking, etc.
A few other ideas are mentioned here by someone who seems to know more about this than I do: https://news.ycombinator.com/item?id=12014243
You rewrote it.
The language you rewrote it in probably has less to do with it than rebuilding all at once what initially was built over time/organically.
That's a very unfair blanket statement.
There are no ways to reduce the inherent flaws of mutable imperative languages no matter how much times you might rewrite the apps in the same language. You definitely can reduce bug surface but not by much; shared data is much more prone to bugs and human mistakes than a system with a runtime that puts immutable data and message-passing front and center.
I'll give you that rewrites do improve state of any random app. But I think you are overestimating this factor with your pretty broad and somewhat dismissive statement.
But in this case the django app was a full rewrite of a php app. It’s quite similar to the eventual elixir app.
I feel it’s pretty much a fair fight, but you’ll have to take my word for it.
The measured workflows in terms of web frameworks include a lot of other elements and none of them are CPU-operations-per-watt (namely pure muscle).
Erlang's (and thus Elixir's) runtime shines particularly well in web apps' typical workflow, not on tasks where C++ likely is still king.
I’ve just shown you that it’s not faster. Benchmark Roda against Phoenix and you’ll see exactly the same is also true for the web workloads.
There is no magic.
(I'll again reiterate that I rewrote a Rails app from scratch to Phoenix and thus accelerated it at least by a factor of 20x. I am just a working guy and have no vested interest in spreading misinformation.)
A lot of web frameworks crumble or introduce abhorrent lags after a certain threshold of pressure. Elixir/Phoenix in comparison almost don't change their average response time until much later.
Truthfully, after reviewing the TechEmpower repo for the last day and a half, I can only see their Fortunes benchmarks as being half-adequate. Most of their benchmarks are extremely unrealistic in terms of web apps usage. At this point I am not even sure what their goal is anymore.
Elixir (Erlang rather) does have some "secret sauce" that makes it performant on part of the web framework space. For example, the literal pool used by the compiler+VM alongside iolists+`writev`-based operations make it so most template renderings in Phoenix allocate as little memory as possible. This means that any literal string in your template is allocated once on boot and not per-rendering, and building the output buffer of the template engine does not concatenate strings either (which would generate even more garbage).
These two blog posts have a bit more detail:
* https://www.bignerdranch.com/blog/elixir-and-io-lists-part-1-building-output-efficiently/
* https://www.bignerdranch.com/blog/elixir-and-io-lists-part-2-io-lists-in-phoenix/
I believe Ruby started to gain some of these benefits with the frozen-string literal annotations. So you can opt-in and no longer allocate static strings on every render. But afaik the template engine still performs a reasonable amount of string concatenation at the static/dynamic boundary. It would also likely be possible to eliminate it by implementing something akin iolists/writev. Maybe it has already been done, I haven't kept up.So in a sense, everything is doable in any of them, but in Erlang this is the default way to do it, so it pushes towards an efficient approach to work with I/O from day 1. Case in point: I didn't invent any of this, I just learned what was there. But you won't see those differences unless you are working on an application that is rendering medium to large-sized templates (which are most apps returning HTML) and benchmarks don't tend to exercise that.
Other benefits from Erlang that may shine in the web space is the per-process garbage collection. It may reduce the variance on the latency and for short-lived requests, you may not perform garbage collection at all. All the VM does is to reclaim the space once the request is over.
At the same time, there are benefits in Ruby and Python you won't find in Elixir. For example, if you need an algorithm that relies heavily in mutability, then Ruby and Python will definitely have an edge. Discord recently had an example of where they made an algorithm much faster by moving part of it to Rust to leverage mutability. But benchmarks don't tend to exercise that either.
In fact, benchmarks may highlight non-ideal behaviour. For fully IO-based exercises, I got faster results by running an acceptor per thread that reads+writes as fast as possible, rather than multiplexing requests on all cores. But in practice, most technologies that can leverage multi-core will prefer to multiplex because that will be beneficial as soon as you do any subsequent I/O or CPU work.
TL;DR - sure Elixir can be 10x faster in some workloads, sure Python can be 10x faster in others. But of course, it would be incorrect to use any of these results to say Elixir or Python is 10x faster than the other.
*Fun fact: I am working on some VM trickery that makes things like Fib/Fac 3-4x faster but at the moment the trickery fails to show any benefit on any slightly more complex code. If they were to accept it, your factorial benchmark would show something much faster for Elixir, but in practice very misleading. :)
This one was a big draw for me. It was a major reason I got interested in Elixir in the first place instead of going the Scala/Akka route.
It's already the standard way in Ruby, even in Rails. Jeremy Evans did the template optimization you're describing using frozen strings 3+ years ago in Erubi and it was adopted by Rails just a few months later.
There's no long STW GC phase during requests in modern CRuby either. It's nowhere near as elegant as BEAM's heap-per-process design but the GC issues seen in CRuby < 2.2 have been gone for almost 5 years now.
GitHub retired OOBGC last year. Pause times are just 1-2ms and total time spent in GC is 0.5% on for the main application I work on.
Are you talking about generating better bytecode based on the semantics of Elixir or a bytecode optimization pass?
Yes, I was aware of the frozen strings optimization. Thanks for confirming! The part I am not fully aware though is in regards to the output buffer, which is a separate optimization. Is the output of a template being stored in an array that is writev-ed to the socket? Or is a template still rendered to a large string? Or maybe something in the middle?
If you don't mind me asking, what is OOBC? My Google-fu failed me. In any case, 0.5% is quite good, although that is also dependent on how much garbage the app generates. And as you know some libs/frameworks tend to be more mindful of that than others. :)
> Are you talking about generating better bytecode based on the semantics of Elixir or a bytecode optimization pass?
Do you mean the VM optimizations I was working on? I was mostly AOT compiling the bytecode of a function. That removes the VM dispatch overhead which further helps with CPU speculation. I see very good results on recursive functions but not much beyond that.
Ahhh sorry I meant Out Of Band GC.
> Do you mean the VM optimizations I was working on? I was mostly AOT compiling the bytecode of a function. That removes the VM dispatch overhead which further helps with CPU speculation. I see very good results on recursive functions but not much beyond that.
It's extremely difficult to get a good return on this approach for two main reasons: instruction dispatch overhead is a very small part of the VM overhead for VMs like BEAM or YARV https://www.sba-research.org/wp-content/uploads/publications... and Haswell onwards have no problems branch predicting the dispatch loop https://hal.inria.fr/hal-01100647/document
Sadly this is also why the CRuby 2.6 JIT isn't very effective either.
V8 removed their baseline JIT because it wasn't worth it. Firefox is considering doing the same. JSC has a more complex but effective baseline JIT. Implementing one for BEAM would be quite a lot of work compared to a minor performance improvement.
foo
<%= expr_that_returns_hello() %>
bar
Imagine it builds its content like this: output_buffer = ""
output_buffer << "foo"
output_buffer << expr_that_returns_hello()
output_buffer << "bar"
With frozen strings, it means "foo" and "bar" won't have to be instantiated on every render, but they are still copied when written to the buffer (unless the buffer uses memory references to frozen strings, akin to a rope).Now imagine the output_buffer is an array, it means less copying, but support needs to be added upstream and other places to handle them. You would also need writev support for arrays (or at least convert them to an iovec or similar).
Does this make sense or am I missing something obvious?
Btw, sorry for all of the questions and thanks for the linked papers. Now I have something to read on the weekend. :)
CRuby makes a bunch of memcpy calls but compared to everything else, they're fast enough they don't seem to matter much.
I've run perf on CRuby running a bunch of different web app type tasks (even rendering huge JSON payloads) and it's never shown up. You can see an example here: https://github.com/k0kubun/railsbench#perf
Perf output for when running Roda looks similar but vm_exec_core and gc_sweep_step are even more of a bottleneck.
This is because sharing reduces the memory usage in more than half. Without sharing (but with frozen strings), the memory allocation is `1 * static_size + 2 * dynamic_size`. I.e. the cost is each static part, which is copied, and twice the dynamic size, which is created and then copied (I am assuming that most of the dynamic content is allocated on the request life-cycle).
With sharing, the cost becomes `0 * static_size + 1 * dynamic_size`. Eliminating the static allocation is particularly important in loops. If you have something like:
<% for ... %>
pre
<%= dynamic %>
post
<% end %>
The factor of `n` applies to both static and dynamic, so you have `n * static_size + 2 * n * dynamic_size` when not sharing but only `n * dynamic_size` when sharing. So you can completely remove a linear component from the equation.You could say that if you use an array or rope, there is still an allocation as the rope or array grows, but that is generally much smaller than the strings being concatenated. It can be preallocated too if you compute the amount of static and dynamic parts when the template is compiled.
So while memcpy is fast enough that it wouldn't show during profiling (except maybe by increased GC times), you should see measurable benefits in terms of memory usage.
Unfortunately I cannot prove this in practice in the context of Ruby, so I will gladly accept if you think I am overplaying the value of the optimization. :)
EDIT: taking a glance at the TruffleRuby rope paper, it seems templating with erb is around 3x faster once they moved to ropes.