Scaling Erlang to run on machines with hundreds of thousands of cores
erlang-solutions.com
erlang-solutions.com
http://www.release-project.eu/
Synopsis of the project:
http://erlang.org/pipermail/erlang-questions/2011-July/06029...
Also, sounds similar to grant to Scala team, Jan 11
host : node : process
So your cluster can have multiple hosts, each running multiple nodes, each node running multiple processes. All is handled transparently by the language syntax -- just sending a message to a process.
I can see the problem being solved on how can these nodes and processes on each host take advantage of multiple cores on that host. Erlang processes are not backed by OS threads in a one-to-one fashion. So how would one allocate these cores to running processes and also how to load balance between 2 nodes for example. One node has 10 processes, the other has 100000 processes. There are 1000 cores available. How to best make use of those cores.
If you have a CPU with 100,000 cores, then you could run the same erlang program on it that you run on a cluster with 100,000 machines with the only changes being addressing the nodes on the cores rather than over the network.
In fact, it is a litmus test for me. When I first started looking at NoSQL solutions, I quickly heard of MongoDB. When I found out it was written in C, I suspected that it wouldn't be distributed and it wouldn't be concurrent... and in researching further, those suspicions came true. (by "distributed" I mean, homogeneously like Riak, not the brittle master-server-sharding setup.)
Trying to write a truly distributed, concurrent application in C is an exercise in frustration-- you'll have to produce halfway done poor implementation of erlang to do it, or you'll deal with pain every single day.
But if you start with erlang, you get these features essentially for free. Plus a couple decades of seriously engineered OTP to go with it.
I know a lot of "developers" these days never use anything more .... intense.... than a scripting language. But if you're ready to dive into the deep end, it isn't replicating the experience of writing assembly in C, it is going towards the future-- erlang.
I, however, would caution against some of the platitudes you use about it being the future of concurrency. I also think it is kind of silly to imply that a coder isn't "worth their salt" if they aren't learning Erlang. It isn't a panacea. If you read some modern Erlang guides like Learn You Some Erlang, for example, you'll note that a responsible advocate of Erlang will caution against drinking too much kool-aid. Erlang, for example, is admitted as being a poor choice for lots of string manipulation because of the way strings are implemented in the language. I believe that same guide mentions that it may not be the best choice for heavy number crunching.
Developers worth their salt will not just hear the word concurrency and start writing everything in Erlang. It is much better to consider the problem and ensure that the problem domain matches Erlang's particular strengths.
I also want to mention that in my experience, Erlang applications are pure hell to deploy for people who don't have years of OTP deployment under their belts. So much of having a manageable Erlang app depends on the developers being well versed in a lot of subtle, and very easy to get wrong conventions. We actually ended up needing to start a support contract with Erlang Solutions themselves to help us demystify getting a sane deployment process going for our app. This was an expensive endeavor. Yes, it was primarily because we were new to the language, but it also had a lot to do with a lack of community documentation/discussion about deployment procedures. I won't even get into the joy of trying to decipher Erlang's famously cryptic crash reports/stack traces.
tl;dr: Erlang isn't concurrency pixie dust. A developer isn't a failure if they aren't currently learning it, nor are they a success if they just decide to use Erlang for every class of problem. It is an interesting language to work with and solves certain problems very well, others not so well.
> Erlang, for example, is admitted as being a poor choice for lots of string manipulation because of the way strings are implemented in the language. I believe that same guide mentions that it may not be the best choice for heavy number crunching.
This are two areas where Go has a big advantage over Erlang while providing an equivalent (although slightly different and more CSP-ish) concurrency model.
> I also want to mention that in my experience, Erlang applications are pure hell to deploy for people who don't have years of OTP deployment under their belts. So much of having a manageable Erlang app depends on the developers being well versed in a lot of subtle, and very easy to get wrong conventions.
Another big contrasts with Go, where you just build a static binary, copy it to any target systems, and run it.
http://www.quora.com/Under-what-conditions-needs-would-Erlan...
I disagree with Erlang "deployment hell". The full OTP deployment process is too heavy for anything non Telecom or Embedded. But for regular Internet applications, you just creating self-containing tar.gz file with OTP release using "rebar generate". Then you upgrade your cluster in round-robin manner. There are some open-source tools to automate this process. But it's easy to use any Fabric-style tool or just your own bash/ssh scripts.
If performance is critical, you can do the string manipulation in C and plug it into erlang fine.
If you're trying to do concurrency in most other languages, then you've probably made a foolish decision, unless those languages are also message passing actor model.
Claiming that deploying erlang is pure hell is FUD, or an example of how you're doing it wrong and really shouldn't be commenting on erlang.
Have a downvote.
There are plenty of good solutions for concurrency. Erlang is one of them. Other message passing actor model solutions are another. Clojure would be another. Using Java's excellent concurrency libraries would be another. Go is reputedly another, although I haven't used it. There is no magic bullet, but there are plenty of good tools of different types.
It helps to 'segregate' your application so it's easier to reason about what data could be a race sometimes, but for me it just made multithreaded applications a little less hard, not easy.
This is grounded in the arguments about how many paths through your program there are; with conventional threading, it's exponential since at any time any thread may reach out and twiddle with something another thread has. Things that isolate threads confine interactions to just their communication points, which is more polynomial than exponential. You can still get yourself in trouble, but at least you don't start out in trouble.
The programming languages we use and the problems we solve with them will influence the way we think about building applications in major ways. In many day-to-day operations, you won't feel the need of powerful string handling. Then at some point you find yourself handling unicode and you find out that the support for unicode is only partial, weirdly documented and somewhat confusing. Then you wish, for that string-heavy problem, that you had a language that handled them better.
Of course you can get Erlang to communicate with other languages, but it's not the same; it's more tedious, comes with its own share of problem.
Erlang is only one of the many available concurrency models: Actors, CSP, Join calculus, models based off pi-calculus, etc. Erlang's model has grown from pragmaticism, but is far from the only good option out there.
Deploying Erlang isn't especially hard (hopefully less than before after re-documenting a few features like releases and relups), although live code upgrades of releases is way too complex for the common project without some solid OTP experience.
When you need save memory and/or CPU when processing strings, you can use binaries, IO-lists and in some cases even atoms. The default string representation as linked-lists is good for 50% of applications, but you shouldn't use it when memory consumption or performance is matter.
e.g.
fold (+) 0 myArray -- return sum of elements in myArray, in parallel
map (*2) myArray -- return a new array with each element in myArray doubled,
-- in parallel.
You can see that `fold` and `map` abstractions that hides underlying parallel execution. Parallelism is transparent to the programmer.My current belief with regards to the future of concurrent programming: concurrency will be hidden behind abstractions from the programmer to the point that the programmer won't (need) be aware of it; Just as when you are programming ruby & python you are not aware of the underlying assembly instructions that are being executed.[1]
Is this what Erlang is like, or, do you still have to deal with race conditions?
[1] Haskell Accelerate is an experiment in this direction: https://github.com/mchakravarty/accelerate/blob/master/accel...
Erlang can be concurrent on a single core, but things can hardly be truly parallel that way.
Erlang's parallelism will thus be coming from being able to parallelize concurrent units, something that usually takes place at a higher level than function application for lists and whatnot (although you can use rpc:pmap for your map example). They're not exactly the same.
Erlang's concurrency is organisational and explicit; the objective is to delimit subcomponents of a system in well established and isolated units, which can then communicate through message passing. The main objective is not to get things to run faster, but to help manage complexity in large systems.