The Erlang Shell
medium.com
medium.com
Those are very tricky indeed, mix in threads with pointers and a system becomes haunted. "This one customer noticed a crash, on this hardware, after running for 10 weeks, but ... we couldn't reproduce it". "Oh I suspect there is a memory overwrite or use after free that corrupted the data". People start doubting their sanity -- "Oh I swear I saw it crash, that one time, or ...did I, maybe I was tired...".
Someone (could have even been Jesper Andersen, the author) said that the biggest performance gain an application gets is when it goes from not working or being crashed to working. And the biggest slowdown is also when it goes from working to crashing unexpectedly.
There was talk of 60 hour weeks here before, one of the things that happens at 8pm at night is people huddled over a keyboard debugging some of these hard to track bugs. Managers and some programmers see it as great heroism, pats on the back for everyone, but, others see it as reaping the fruits of previous bad decisions and it is a rather sad sight for some. It all depends on the perspective.
I guess the point is one of the main qualities of Erlang is not concurrency but fault tolerance. Many systems copy the actor concurrency patterns and claim "we have Erlang now but it is also faster!", and that is a good thing, but I haven't heard too many claim "We copied Erlang's fault tolerance, and it is even better than Erlang's!". [+]
[+] you can do the same pattern up to a point using OS processes, LXC containers, separate servers, having watchdog processes restart workers, for example.
A concurrency idiom I've gotten a lot of mileage out of is to use threads, without shared state, but with pointers. Instead of shared state, you use a message passing style where you pass pointers over queues. Threads only mutate data that they are explicitly handed via a queue. After putting an object on a queue, they NULL out their local reference.
It's basically like Go's goroutine idiom, but with C/C++/Python, real OS threads, and a thread safe queue (which is all a channel is). There is also some relation to the SEDA architecture.
The logical structure of such a program is not really that different than an Erlang program, I imagine. It just works a little better with existing C/C++ codebases, and is potentially faster in some cases.
It has a completely different character than spaghetti written using threads and shared state, and can be made extremely reliable. Especially because the explicit channels provide a hook for testing black box behavior.
It is a continuum. One could write it with shared mutable state too, using locks. The problem is it is easy to mess up.
C, a large program could load modules, and some of them are thread safe some are not. One could bring in a new module that calls some initialization routine (say curl_global_init()) and then another part of the system ends up calling it as well. But it is not thread safe and should only be called once. It is not intentional sharing, but it just happens as it is running in the same memory space.
Someone likened that to running your code in Windows 3.1 when you Command & Conquer game would crash and take down the word processor with it. They didn't intend on doing so but it happened by accident as a bug. Sometimes running your application as a Windows 3.1 system is acceptable, sometimes it isn't. It depends on the problem.
If the library has non-reentrant code, then you can't have more than one thread for the stage -- you can't parallelize it. But it generally won't lead to correctness problems.
curl is actually a great example, because it has an event loop for parallelism (rather than threads). So you would run a single thread for curl, because you wouldn't want more than one thread anyway.
Say you are writing a web crawler. From the network/curl stage, you just pass off pointers to blocks of memory to parsing threads. Parsing threads will be CPU bound so you will likely want to run instances of that loop in multiple threads, and you will be able to do it with no problem, since they don't depend on the curl library. They just take in blocks of memory and output some data structure to another queue.
This is also a good way to compose say the curl event loop with event loops from other libraries (GUI libraries, perhaps). Hence the relation to SEDA (http://en.wikipedia.org/wiki/Staged_event-driven_architectur...).
I can hook up to the Erlang runtime (in local, staging, production, wherever) and talk to these actors. I can ask them "hey, what's your state now," "who are you talking to?" Sure, Erlang is slower in some cases, but for the value you get, I think it is priceless in many cases.
I am a big fan of interpreted languages and use them as my default. The situation I was describing was basically the only time I've ever rewritten in C++ for speed! It just doesn't happen that often.
But I do wish the interpreted languages like Python, node.js, and Erlang all had better and more consistent C APIs (more like Lua's). C itself is not that hard, but the C APIs definitely put people off.
Actually I wonder for Erlang, with the web crawler example, would you have to copy entire web pages if you wanted to pass them off from a "network actor" to a "parsing actor"? The threads + pointers solution easily avoids that, while retaining modularity.
Also, when you split, or reference sub binaries, these become just pointers into the original binary data.
If you don't have a queue library and don't want to risk getting the locking wrong, you can get one from pipe(2). While simultaneous writes/reads are in theory permitted to be interleaved, the implementation would have to be actively insane to interleave anywhere other than page boundaries.
...I actually did this when I was at university, for the first class that covered locking and concurrency. Amazingly enough, I didn't fail that assignment for delegating what we were supposed to be learning to the OS.
The logical structure of such a program is not really that different than an Erlang program, I imagine. It just works a little better with existing C/C++ codebases, and is potentially faster in some cases.
It has a completely different character than spaghetti written using threads and shared state, and can be made extremely reliable.
Maybe this is why I don't get all the C++-hate? The things that make code unreliable or thread-unsafe tend to also make it hard to read and think about (even without threads), so I just find some other way that doesn't give me a headache.
Interestingly code that doesn't give me a headache tends to have a strong resemblance to functional-style code, even tho actual pure functional code also tends to be mildly headache-inducing.
Yeah, there are a million ways to use C++. To be fair, most people have their architecture forced upon them, rather than creating it from scratch. And the C++ world does have a culture of spaghetti with shared mutable state and threads. In Erlang you have the opposite culture. So I can understand the association that people make with a language, even though it's more a matter of culture.
In addition, C++ doesn't even have a thread safe queue in the STL (although I was pointed to a draft for one in the next version of C++ by a coworker). pipe() is a good solution I would use, but not portable to other OSes.
I come from a Rails background. While Erlang's hot code reloading and the ability to attach to running nodes is cool, I am not quite able to connect the dots on how this is useful. In Rails, a process is short-lived by design (the lifecycle of a request). Attaching to a process is useless, but the Rails console essentially gives you an isolated process to inspect production code and data. I don't see how extending that to an existing process would be useful.
So, challenging my assumptions, perhaps short-lived processes aren't the best approach for a web service. Perhaps that's the paradigm in Rails because Ruby lacks the tooling to make a long-lived process easy to maintain. If a long-lived process is as easy to maintain as short-lived processes, then suddenly new architectures open up. How are these architectures different? What are the benefits? What are some examples?
Simulating something that is designed to work with thousands or even millions of users at the same time is pretty much impossible. Hot code loading and attaching to console gives you the option to see what is going on in real time and even fix it if you have to.
Now you should probably have a supervisor. Your connections could be handled by a connection or socket pool. You could have a database driver. There could be a background processing queue. Now things get interesting, not all those things get torn down and recreated on each request.
Even for short lived requests you could attached to a live system and trace one of the short lived requests or set a debug point to see what it does in "slow motion" to so speak.
One can use Erlang for quite a bit without touching the distribution (as in distributed nodes across many systems forming a single cluster). Then the shell can help you log in to any other node, transparently.
The ability to reload code at runtime is not something many systems can boast. Because that is one case you might touch once you have found the problem, but that is again one thing you can do.
You can diagnose loading and system resource utilization as well. If you have a fully-featured system (say downloaded from Erlang solutions) open the shell and type observer:start(). to start the observers. It will launch a GUI window (make sure to have X forwarding if running via SSH). There you can see runtime stats, the application tree, process list, ETS tables. For each application in the tree you can click on the process and see its state.
In Erlang, you would typically answer a web request in a new actor every time. Just as in Rails, running for example N number of threads serving one request each. Of course, in Erlang that number of actors can vary, which means your system will be more responsive and you'll have lower average latency. However, what differs mostly, is that in Erlang all those actors live inside one VM and you can have central actors managing things, collecting statistics, keeping global state, talking to databases / other services. This, I think, is the core strength of Erlang compared conventional single threaded / GIL systems.
A more concrete example would be to dynamically show the current request rate on a web page. You can just create an actor that gets a message from any other actor once they handle a request, and keeps track of how many such messages per second arrive. Then when rendering that web page, you just ask that actor for it's current value. Simple, beautiful, dynamic. No need for a database or any other central storage.
I'm not sure I agree. There's https://github.com/ChicagoBoss/ChicagoBoss and quite a few other frameworks (https://github.com/ChicagoBoss/ChicagoBoss/wiki/Comparison-o...) and there's also Elixir and some frameworks for it (https://github.com/dynamo/dynamo).
It's of course nowhere near the ease of development of simple web pages in Django or Rails, but it's not like you need to reinvent everything from scratch either. And you gain ability to scale almost without any effort. It's a tradeoff, of course, and using Erlang for every web page you build from now on would be probably an overkill. But there are web sites which are not services yet, which could still benefit from Erlang greatly. And I think (and hope) that the initial overhead of Erlang will get smaller and smaller over time as new frameworks emerge and that someday deploying your blog on top of Erlang will become viable choice :)
No, not everything, but lots of stuff.
Recent example: I had to fix the Postgres database driver to better handle queries like
"select * from foobar where id = any($1)",
[uuid1, uuid2, uuid3]
Also, ChicagoBoss is a bit rudderless at this point in time - Evan Miller has moved on to other things, and the guy who briefly took his place has as well. My own branch of it on github seems to be the only one getting any attention lately, and that's just small fixes here and there:Yes and no. Rails processes are so huge that they are loaded up and kept around between requests. Erlang processes (which are not unix processes!) are much smaller and lighter. One Erlang program can handle tons of concurrent, always open requests, whereas Rails just isn't cut out for that.
As to what Erlang is good for, I think its sweet spot with web stuff is where you need a lot of always open connections, such as with web sockets.
I'm a fan of Erlang, but for most SaaS kinds of startuppy things, am not sure it's a good fit: you get so many more tools packaged up and easy to use for you with Rails that unless you know Rails does not work for you, I would almost always choose Rails.
http://learnyousomeerlang.com/
Free on the web, but buy a copy - it's worth it.
I'm not sure if this is by design, or a limitation of Ruby.
>short-lived processes aren't the best approach for a web service.
Not for me - I'm building all of my web services in Erlang these days, and there's little to no relationship between the life of most processes in the service and the life of requests. Although each request is usually serviced by its own process, that process may interact with other processes that may have been alive for as long as the application has been running. For example, my RAM caching strategies are usually written as part of the application, not put in front of them. Makes pseudo-ESI easy to hand-roll.
Also, short lived processes may have been spawned by long lived processes, so changing how they're created can be done without taking anything down.
I find this style so much more performant than the one thread per request thing from Ruby, Perl, PHP et al. It's very easy and natural to minimize the amount of time spent in the db (which is where I'm usually bound.)
A few questions and comments:
1. RAM caching - I can see a big benefit to keeping cache in Erlang. No marshaling of data to and from memcached, for example. How do you handle eviction?
2. Avoiding the db - I'm going to guess this ties in with RAM caching. Are you doing a write through cache or something like that?
Thanks for your response, it's helping it all to come together for me.
I am guilty of this. I am relatively new to erlang programming and I have skipped on some of the features that make the language great. For instance I have neglected learning both supervision and hot code reloading.
The biggest thing to get on your initial survey should be minimal awareness so you know that that feature exists when you encounter the problem it was originally built to solve.
Supervision though, is fundamental to OTP and Erlang. It's certainly easy to learn so I would consider that a low hanging fruit for you.
Hot code reloading is a consequence of the continuation passing style that is prevalent in every Erlang process and easy to understand too (although employing it in production can be a bit trickier).
Ignoring supervision, however, is...not wise, if you have -any- sort of reliability considerations (even ones as weak as "I'd prefer it to remain up"; if it's just a one off thing that you'll manually execute again if it fails, meh, fine).
Supervision is really, really easy to use, at least for the base use case (creating a good hierarchy, so that processes restart the other processes that should be restarted, is an architectural consideration that can be a bit hairy), and it's truly heartening to grep a log for something you've been running a while, find crashes that indicate issues you want to fix (even simple stuff, like sanitizing inputs), but the system is still running as though nothing happened.
GDB or even VS's debugger can do a lot of magic even if you use C. Until absolutely necessary, leave debugging information reasonably high and optimization reasonably low or off and you'll get an amazing, dynamic environment to inspect a live C system.
Yes.
> I'm genuinely curious, I kind of had this impression that it was a feature fairly unique to Erlang.
It's not. iIRC there's even a version of the JVM with code reloading.
What's interesting about Common Lisp is that not only can code be dynamically compiled into the running image but class definitions can be changed and live instances will be updated without stopping the program or reloading anything. There are a tonne of features in Common Lisp like this that are geared towards robust, maintainable software.
Yes, because of Java's extensible classloading mechanism, all you have to do is drop an war/ear to the 'deploy' folder of your application server. New files are deployed, rewritten files are redeployed and deleted files are undeployed.
One of the beauties of Erlang's hot code loading is the ability to migrate internal application state. An example:
I'm storing a lookup table going from key to value as a tuple {key, value}. With a change going out, I'm going to start doing {key, {value, other-value}}. In Erlang there's an easy pattern in place to do that. Just provide a function that goes from {key, value} to {key, {value, sensible-default}} and as part of the code reloading process that function will be run before the new code starts executing.
That's not at all a perfect explanation of the process, but it's more or less useful enough to describe the difference between the two platforms.
if you need more elaborate lifecycle management of components as they are upgraded, then maybe OSGI could do that.
erlang's approach really appeals to me though -- not only does the code get upgraded, but it's in cooperation with its message loop so you don't miss a step, and you also get a callback to migrate your actor state.
I've seen some info about the JVM implementation, but I could've sworn it wasn't able to preserve state while hot swapping. I haven't dug in to it though, so it's totally possible it handles that.
> There are a tonne of features in Common Lisp like this that are geared towards robust, maintainable software.
I'm a fan of Lisp in general, though my experience is limited to a tiny bit of Scheme and experimenting with Clojure. I'm glad to see that stuff like this is available in at least one of the Lisp variants! Question for you (if you program in Common Lisp): What are the main selling points? I spend a lot of time already with a couple of languages, Elixir/Erlang (for fun), C#/F# (for work) - what is it that would make me want to switch from one of those to Common Lisp?
Thanks!
Erlang has them in the form of parse transforms and Elixir has them too and makes them easy to use. Not as easy as sexp based languages, but almost.
1. Configurable compiler. You can tell the compiler where it's safe to ignore types and where to be strict about types. For the battle-hardened critical sections of code turning off run-time type checking can be a pretty big win performance-wise.
2. Extensible compiler. Macros are like tiny compilers. It's possible to create DSLs for the programmer to work with a problem in terms of nouns and verbs that make sense for their task while still compiling to VHDL for your application. Common Lisp becomes the substrate on which you build the language you need to solve your problem.
3. The best error handling system I've seen in any language. Signals, handlers, and restarts allow higher-level code to choose an appropriate strategy for recovering from errors under various conditions without losing state and restarting the computation.
4. Top-class introspection tools. The JVM has been able to do some of this stuff for a while but with CLOS and the error handling facilities available Common Lisp makes this so much better. You can inspect any object, change values on slots, trace any function, step through the entire stack, and recompile code into the running image. There are nice tools written on top of these facilities which provide some of the most elegant debugging sessions I have ever encountered. Stop your program? Why? I can connect to a running image, inspect the running program interactively, and compile the fix into it without restarting. Even over SSH directly from my IDE.
And that's just the stuff in the specification. I'm a fan of the newer ASDF systems for defining systems and their dependencies (implicit file-system structure for defining module composition? No thank you). There are some really solid libraries available (iolib, bordeaux-threads) and a very nice FFI for interfacing with alien code. And various implementations provide interesting extensions depending on what you need. Clozure Common Lisp, for example, has really good native Cocoa bindings. SBCL has a very slick compiler for generating fast code. Most of them provide threading extensions and there are libraries which abstract away the differences if you're writing a library that depends on threads.
For example, how well CL's class redefintion and instance reinitialization work with preemptive threads? It isn't obvious how each implementation address this issue by skimming docs, so I'm curious if anybody with direct experience. When I implemented CLOS/MOP style OO in my Scheme, I ended up using a global lock during class redefinition, for it need to modify lots of places. It's MT-safe now, but I'm still wary to run class redefinition on a system written in my Scheme, which is busy handling requests.