Like what ?
Like what ?
In my mind, Erlang excels in a few ways:
Less often remarked, binary parsing in Erlang is rather nice. Parsing out an IPv4 packet looks like this:
<<4:4, IHL:4, _DSCP:6, _ECN:2, _TotalLength:16,
_Id:16, 0:1, _DF:1, MF:1, FragmentOffset:13,
_TTL:8, Protocol:8, _IPChksum:16,
SrcIP:4/binary, DestIP:4/binary, IPRest/binary>> = Payload,
OptsSize = (IHL - 5) * 4,
<<_IPOptions:OptsSize/binary, IPPayload/binary>> = IPRest,
And then you've got everything. As long as formats are reasonably documented, it's easy to parse them. Sometimes, you can even parse a size and use the size in the same match.More often remarked; because of the language constraints, most notably a lack of shared memory between processes and immutable variables, most programming ideas end up expressed in a way that's amenable to massive concurrency, while being comprehensible. Most Actors have easy to understand behavior --- they may accept messages, leading to them sending messages and changing their internal state. From there, you may need to puzzle out the overall system behavior, but often times, getting each individual Actor's behavior correct, leads to correct (if hard to verify) system behavior.
Hot loading code reduces deployment time, which increases developer productivity. When you've got a million users connected to a machine, it takes a lot of time to move them to other machines so you can do a traditional stop / start cycle; hot loading means you can fix things in seconds (and, of course, it means you can break things in seconds too). You can, of course, hotload in C, and probably other languages, but very few people do it. It's comparable to pushing PHP files though.
And, of course, the most important thing is ejabberd is in Erlang, and it looks like it has what we need for a chat server, and I heard some other people scaled it really far. ;)
It's because it's not super easy to properly architect in, it's hard to test, and it requires supporting running multiple versions of the code concurrently (and the ability to migrate data on the fly).
Keep in mind that we rolled out the patch incrementally to one or two machines and then to more. If it looked good we’d roll it out to the rest of the cluster. From there we’d make note of the result and get the change ported into a proper release.
Each time I hear stories that both clustering and hot upgrades, module loading, etc are hard or rarely used I wonder if the person is just repeating someone else’s rumor. They’re great features and work fine if you do a little homework (just like learning anything, don’t treat it like magic).
This was from my time at Cloudant/IBM, though its far from the only case I’ve seen. We ran over 1000 machines this way with some clusters growing to more than 200 nodes using distributed Erlang (something I keep hearing is hard or impossible, it’s not).
Badfun errors are a great example of something tricky if you're not used to thinking about how the compiler translates these to private functions via lambda lifting. It's the same reason funs are not something you should rush to pass around over a disterl cluster. Recursive functions which keep fun's around in a loop also become a problem if they're long lived.
I usually point out that first class modules or MFA tuples are more idiomatic in Erlang than opaque first class functions as values but we're off into the weeds here. It's a good example of where effort and gotchas become a barrier for many.
Lots of languages have actor systems. But how many have preemptive multi-tasking on these actor systems? I can't think of any at the moment. I am well versed in Akka for Scala and the lack of preemptive multi-tasking for actors is a big pain in the ass. Erlang's advanced BEAM and support for this is it's main selling point, imo.
You have a very naive view of large-scale distributed systems. For companies that truly need them, "time" and "effort" are never a consideration. Furthermore, the challenges facing these companies in battling project delays and cost overruns are the 100% organizational and political, not technical.
My entire intention was to say Erlang is easy to get off the ground with quickly.
Most companies outside of the Silicon valley bubble are resource limited.
a) Erlang (as opposed to something else) is only useful for large distributed systems.
b) If your organization is at the scale where you need a large distributed system, then the problems in your project aren't related to code or to coding speed.
b) Needing a large distributed system is a requirement that one can get to from lots of situations. You are making way too large of a generalization.
b) No. Large distributed systems is not something you stumble onto, it's a result of many years of organic growth.
I'm not sure that's true. Sure, a few of the most famous software companies in the world have built their own at considerable cost. But numerous other huge, important companies are dependent on large-scale distributed systems whose major drawbacks (reliability and maximum scale, usually) and major benefits (simplicity, time/resources saved on not having to hire tons of specialists, quick development time) are based precisely in the time and effort constraints under which those systems were developed.
Probably because that claim makes no sense.
Do not assume you'll be able to build your own Whatsapp just because of Erlang.
The story is greatly underreported and all focus is only on "they run Whatsapp on Erlang with just ~50 engineers".
Highscalability lists just some of the patches and optimisations they have here [1] and here [2]
Here's an incomplete list of patches only. There's also tuning and optimisation:
Erlang: Fixed head-of-line blocking in async file IO by patching BEAM, added round-robin scheduling for async file IO, added multiple instrumentation patches. Instrumented scheduler to get utilization information, statistics for message queues, number of sleeps, send rates, message counts, etc. Made lock counting work for larger async thread counts. Patched to dial down spin counts so the scheduler wouldn’t spin.
BSD: Backported a TSE time counter. Backported igp network driver.
More Mnesia (Erlang) patches discussed here: [3]
Are you ready to do this for your Whatsapp?
> elixir/phoenix also achieved the same "2 million connections on single server" without needing those optimizations.
There's more needed to run a chat server than just "2 million empty connections".
---
[1] http://highscalability.com/blog/2014/2/26/the-whatsapp-archi...
[2] http://highscalability.com/blog/2014/3/31/how-whatsapp-grew-...
[3] https://www.infoq.com/presentations/whatsapp-scalability/
I don't need to thanks to the fact that a bunch of those patches are now part of Erlang.
There's also signinficant tuning and optimisation, both for the Erlang VM and FreeBSD.
There also things like (quotes from Highscalability):
"Mnesia: Using no transactions, but with remote replication ran into a backlog. Parallelized replication for each table to increase throughput."
"When Rick is going through all the changes that he made to get to 2 million connections a server it was mind numbing. Notice the immense amount of work that went into writing tools, running tests, backporting code, adding gobs of instrumentation to nearly every level of the stack, tuning the system, looking at traces, mucking with very low level details and just trying to understand everything. That’s what it takes to remove the bottlenecks in order to increase performance and scalability to extreme levels."
Or even the things like "What has hundreds of nodes, thousands of cores, hundreds of terabytes of RAM? The Erlang/FreeBSD-based server infrastructure at WhatsApp". Oh, wait. Erlang's default distribution mechanism grinds to a halt when there are more than ~60-80 nodes. And Mnesia has a 2GB limit on table sizes. So you have to work around those limitations yourself.
There are no magic bullets. Erlang will only take you so far. The rest (80-90% of the way) you have to take on your own, and you have to know what you're doing, and what needs to be done: patches, tuning, workarounds, limits of the systems you work with etc.
I've seen people say these, and I have no idea where they come from. If you have a decent network, dist works fine at well over 80 nodes, but everyone says it doesn't work. pg2/global has some sharp edges if you're trying to have many nodes acquire the same global lock when you have a lot of nodes (a few hundred) or a smaller number if you have a lot of latency between them. There's options though -- maybe you don't need to acquire the same lock on all nodes, or maybe you can look in pg2.erl and global.erl and wiggle the locking code until it no longer live locks.
The Mnesia supposed 2GB limit is a bunch of hooey. Yes, disc_only_tables has (or had) that limit, because dets has that limit. Yes, it's a sharp edge, because there's no warning about it. However, a 2GB dets table is awful to work with anyway. You want to use disc_copies or ram_copies for big tables. Also, mnesia_frag is well supported, so if you really wanted to, you could make your disc_only_copies table 1024 fragments, and have 2 TB of dets, if that's how you wanted to role.
And yes, if you're going to hyperscale, you're going to need a couple people who know how to figure out what your system is doing. Is there a language/environment where that's not true?
I claim, without real proof, that Erlang's BEAM VM and OTP standard library are easier to understand and tweak when you do hit problems. You'll note however, that Rick Reed's first presentation was when he had been at WhatsApp for about a year, and he had zero experience with Erlang before that.