[1] https://www.techempower.com/benchmarks/#section=data-r18&hw=...
[2] https://elixirforum.com/t/techempower-benchmarks/171
EDIT: Thanks everyone for the responses.
[1] https://www.techempower.com/benchmarks/#section=data-r18&hw=...
[2] https://elixirforum.com/t/techempower-benchmarks/171
EDIT: Thanks everyone for the responses.
For a plain-text benchmark, you are ultimately measuring the ability of sitting on top of a socket and reading/writing to that socket as fast as you can.
However, Phoenix/Cowboy, whenever there is an HTTP connection, spawns a VM light-weight process to own that connection, and then each individual request runs in its own light-weight process too. This comes with its own guarantees in terms of state isolation, reasoning about failures, and you get both I/O and CPU concurrency for free - if you were to do any meaningful work on those requests.
Even though the VM processes are cheap to spawn and lightweight, it is overhead compared to something that is just directly reading and writing to the socket. I actually wrote a proof of concept where we just sit on top of the socket without spawning processes and it performs quite well - although I don't think it has any practical purpose. There is also an interesting article from [StressGrid](https://stressgrid.com/blog/cowboy_performance_part_2/) that shows how removing the per-request process speeds it up by ~60%. But once you are going through long requests (i.e. 1ms because you need to talk to a database or API, encode JSON, etc), these differences tend to matter less and the concurrency model starts to give you an upper hand.
> But once you are going through long requests (i.e. 1ms because you need to talk to a database or API, encode JSON, etc), these differences tend to matter less and the concurrency model starts to give you an upper hand.
If you look at the other tabs in the benchmark there are "heavier" loads that include database access and the most favourable one I can find is "Fortunes" which puts Phoenix at 17.8% of the relative performance of the best ASP.NET Core one.
On the plain text one, the difference between frameworks is at most 6x. You would have to yank Phoenix and sit directly on top of a socket, as per my previous comment, if you want to compare both VMs.
(This is not to say you can’t tune BEAM to the needs of other workloads; those are just the defaults. Though, because that is the default assumed workload, it’s also the one most optimization effort goes into.)
BEAM is also pure-functional in a way that a lot of VMs hosting pure-functional languages aren’t; things like mutable atomic memory handles were only introduced recently, and there’s no plan to rewrite core operations in terms of them, because these primitives are less predictable in their time costs than equivalent functional primitives.
You can go fast under BEAM, but the people who want to do so, often decide that it’s too hard to go fast while also retaining soft-real-time guarantees they began using a BEAM language to attain; so they instead use one of BEAM’s many IPC bridges to hook up a fast native “port program” written in another language to the BEAM node, ensuring that any hiccups or crashes in the native code are isolated to its own address space and execution threads.
This last consideration means that very few people are really driving BEAM’s performance forward, because so many instead take the “escape hatch” of low-overhead IPC.
Taken as whole systems, though, you’ll find BEAM in a lot of services that go very fast indeed—you just won’t usually find BEAM itself handling the performance-critical part!
> It's incredibly good at multiprocessing, which is really useful in the websphere
So on that benchmark I linked, if you switched to the latency tab Phoenix clocks in at 14.6ms average latency with a standard deviation of 23ms and a max of 460.7ms. The best looking ASP.NET Core result from a standard deviation and max point of view is 1.5ms average latency, 1.4ms standard deviation and 49ms max.
So again it's hard for me to tell whether it's BEAM or Phoenix that's responsible for the relative performance hit.
There’s also the fact that Elixir and even Erlang don’t take full advantage of the possibilities of BEAM-bytecode whole-program optimization. For example, any list you walk using Erlang’s lists module or Elixir’s Enum module is going to generate a back-and-forth of remote (symbolic) calls, rather than resulting in the body of that function of lists/Enum being inlined and fused into the caller. Which in turn means, even if you use HiPE, you’ll be thunking in and out of native code so much that the overhead will eat any potential benefit; and each side will be opaque to the static analysis of the other, so even a wholesale native recompilation of the runtime won’t save you. (Same with gen_servers: remote callbacks through library code, rather than fused module-local receive statements; ergo, 10x optimization opportunity lost.)
You’ll see real optimization in code generated by e.g. Erlang’s lexer-generator library leex; or in specific parts of each language’s stdlib, like Elixir’s Unicode-handling functions. But such optimized code emission is thin-on-the-ground compared to “naive” coding in Erlang/Elixir. (And I don’t blame the language devs: the languages are fast enough for almost everything people try to do; and for the rest, you can ignore the language-as-framework and write entirely self-contained modules that only rely on BEAM primitives.)
I thought it's exactly the other way around.
I Always assumed BEAM is built for concurrent, consistent low latency IO. The benchmarks seem to vaguely reflect that.
1. that an actor that wants to run is never de-scheduled for too long (= 99th percentile latency);
2. that every actor that wants to issue an IOP to the kernel gets to do so efficiently (= whole-system throughput)†;
3. that any individual actor makes actual appreciable progress (= 50th percentile TTFB latency.)
This is basically the definition of a soft-realtime system. If latency always came first, it'd be a hard-realtime system; if throughput always came first, it wouldn't be a realtime system at all.
† BEAM is actually quite bad about this for disk IO in the normal case, because all disk IO by default goes through the bottleneck of sending messages to a single file-server actor. This has the advantage of virtualizing the filesystem, allowing for e.g. diskless BEAM nodes that boot by using a peer's file server as their own. But it's not very good for filling up your disk's IO queue. Luckily, you can bypass the file server and use file-IO ports/BIFs directly. Or, if you're working against network storage anyway, just skip your own OS kernel and have BEAM speak iSCSI to the disk server directly (like Erlang-on-Xen does.) That approach will likely give you more disk IOPS than you could achieve even in C, presuming the C program is relying on OS syscalls to talk to the remote disk through the OS VFS layer.
However, I meant that I always assumed that BEAM prioritized low consistent latency at the expense of high throughout (but without providing hard real-time guarantees). And the mechanism being that preemptive scheduling and counting those reductions so no single process is able to hog the entire system's resources for expensive computations.
While these benchmarks may not be reflective of the true performance it seems that Phoenix has a smaller throughput but lower latencies than most of its peer group in them.
One actor queuing up a heavy amount of IOPS through many different ports, can actually cause the BEAM node to generate a visible latency spike, as it always tries to schedule all the ready ports assigned to the given OS scheduler-thread per scheduler-timeslice, before trying to schedule any user-defined actors; and so will find itself with very little time left in that timeslice to schedule the execution threads of actual actors. A single actor can accomplish this “uneven” loading because BEAM’s reduction scheduling means that just queuing IO ops through ports in a linear block of code can never cause an actor to be descheduled, since the port BIFs are not function calls per se, and therefore have no ability to trigger scheduler yield-checking.
Mind you, this is a highly-unusual usage pattern. Actors don’t usually batch up huge binary payloads to send and then blast them all through at once into many different ports in an unrolled loop. And, usually, one actor doesn’t even usually hold multiple ports; the idiomatic approach is for each port to have its own separate “owner” actor (because this means that these “owner” actors will be scheduled semi-synchronously in response to activity on their respective ports, sort of like Go with goroutines getting scheduled in response to activity on channels.) In an idiomatic usage pattern, you could say that the node doesn’t become IOPS-bound but rather actor-scheduling-bound, where a scheduler-saturating number of ready ports can’t be brought about, because as ready ports per timeslice increase, the amount of the scheduler-timeslice available to actors decreases, and so fewer actors that would generate IOPS are getting run, meaning that more of the next timeslice will be available. In other words, with idiomatic use of ports, this is a self-limiting problem.
But, well, we’re talking about how the BEAM is built, rather than about what Erlang/OTP systems do in practice. The BEAM itself has these priorities, even if, in practice, they’re papered over with other ones :)
Elixir is pretty slow for simple singlethreaded CPU bound tasks.
It's incredibly good at multiprocessing, which is really useful in the websphere since your workload is much more likely to resemble thousands of simultaneous communications that need to not block each other rather than calculating a 13042345235223452 digit prime number.
Right so the benchmark I linked to was TechEmpower Web Framework Benchmarks, and Phoenix managed 186,774 requests/sec on a 14C/28T Xeon to do basically nothing (echo Hello World) and ASP.NET Core managed ... seven million.
Plus, 186k/sec is still a hell of a lot.
https://medium.com/@alexhultman/millions-of-active-websocket...
The answer you are looking for is: much slower than .NET core.
The "being good at multiprocessing" of the BEAM VM people are bringing up is not about raw speed at it: it is about giving you the tools to actually being able to do multiprocessing in a sensible, stable manner.
For instance, what happens if a bug in your program blocks threads for a long time in .NET core? What if one thread crashes? Do you have any guarantees that one thread cannot affect others in any way? Etc.
Some developers think that giving up raw performance for guarantees in these areas makes sense. Other developers think they can deal with those issues without nice guarantees from the platform and thus prefer the speed. There's no definitely right answer...
See the architecture of Phoenix Channels (which powers LiveView and other websocket-based solutions in Phoenix): https://hexdocs.pm/phoenix/channels.html
To expand on your workload description: things like waiting for a database query or an API result are smoothly handled in this paradigm; the process goes to sleep for a bit whle another request is processed.
If your workflow is like a delivery truck which makes a lot of stops, "faster trucks" doesn't help much with throughput. "More trucks" is better.
The aspcore platform benchmark does some cool trickery based on whether the test even requires a database using compiler #if directives. Though it seems to be flagged as "realistic approach" Dunno if I agree with that. I think I'm looking at the right code. https://github.com/TechEmpower/FrameworkBenchmarks/blob/mast...