That's actor model, but it still doesn't feel as natural as with functional languages - dealing with immutable data in OO languages is tedious and error prone
That's actor model, but it still doesn't feel as natural as with functional languages - dealing with immutable data in OO languages is tedious and error prone
aka the original OO. OO as taught today is nothing like how Alan Kay imagined OO [1]. Alan Kay originally imagined OO as biological cells only able to communicate via message-passing. That's the actor model more or less.
> dealing with immutable data in OO languages is tedious and error prone
Again, pretty different OO. Still, I'm curious what you find error prone? Tedious, perhaps. Scala and Java+Lombok make immutable data pretty decent to deal with ergonomically if not quite as good as Clojure or other languages out of the box. Even PCollections for Java reduces much of the waste and is fairly nice ergonomically.
[1] http://userpage.fu-berlin.de/~ram/pub/pub_jf47ht81Ht/doc_kay...
Async message-passing has a huge problem when it comes to practicality: In any resource-constrained system (i.e. no unbounded queues) messages can get lost. I don't know if you feel comfortable with method calls having a "best-effort" semantic (which is what happens in the actor model and biology[2]), but it makes me deeply uncomfortable as a foundational model. At the very least you need some way to do retries, but if even looping is a method call, then how do you ensure even the reliability of retries? You could do it to arbitrary precision, I think, given enough duplication + repetition of code[1], but I don't understand why we'd want to do something so messy at a language level when we already have hardware that's reliable to 99.999%+ levels already.
[1] This also assumes that the interpreter/compiler reading the code is 100% reliable... which it wouldn't be if everything is actor-based. You can see the problem.
EDIT: [2] This is a perhaps-interesting aside: In biology "best effort" is typically the best you can do, given the limitations of fluids/energy/etc. You send out enough molecules of the right type to hope that the intended recipient has a decent chance of detecting them. These things have been fine-tuned through a billion+ years of evolution, but in computer systems we typically don't have that kind of (even simulated) time to do "just enough".
In fact the prevalent view is that code is unreliable and you should structure things to recover from failure at all points of the architecture. Messages can go missing. Calls can time out. Processes can crash for reasons you will never know about, and your architecture should be resilient against that.
Whilst this sounds extremely complex, in practice it boils down to letting everything low level fail, and having supervisors who can handle the failure.
Resource issues are a problem everywhere. For example, you want to save a file.
What happens if the file system crashes after it claims it has saved?
What happens if its a network connected drive, and it times out?
What happens if its a file on Amazon S3, and your network cable just got cut by a tractor?
You're going to have to deal with this at a higher level than the line of code which tried to save the file.
Exactly, but that's the point I'm trying to make, perhaps badly. My point is: Can a function call fail in Erlang, i.e. can it just "not return"? AFAIUI as long as you're inside a message handler, this cannot happen -- any function call will receive a return value. In other words: Function calls inside message handlers are synchronous.
In the Actor model conceived by Hewitt, this is not the case. Now, this would be a huge problem for any foundational theory, but he seems to get around it by just assuming infinte resources (queues), so it all works out in theory.
Kinda? Function calls in Erlang are intra-process and synchronous, but the function could send and receive messages internally and never return because it never got a response and did not have a timeout setup.
When you're talking about actual messages between processes though, they are asynchronous and they can get lost (if only because the process you're sending a message to is on an other physical machine).
Regarding applying backpressure, you're correct that Erlang doesn't have a silver bullet. Process mailboxes are unbounded and can exhaust resources. Or you can implement a buffer and drop messages based on some strategy.
As for making every 'operation' asynchronous, you could do it and not run into any additional unbounded queue or error handling problems, but it would add overhead without any advantage over concurrency via preemption.
- Synchronous calls where the caller will receive a response, and waits for X seconds until considering the message lost. Synchronous messages are blocking on the caller side, and block on the receiever while they are being processed.
- Asynchronous calls where the caller fires the message off and further semantics are undefined---the message could be ignored, lost, received, and acknowledged through a side channel, or a message back to the caller later on.
The Erlang runtime provides a framework--OTP as in online telecom protocol referencing its always on, phones must work at all times roots--that lets you set up supervisors, watchers, and handlers to deal with the problem of trying to send messages to processes that aren't there, child processes disappearing on you, and so forth.
Erlang's ecosystem aims for a 'soft realtime' setup, where usefulness degrades but is not totally lost---so the default option on failure is to retry until some kind of success threshold is reached.
In Erlang/OTP it's wrapped in an actor behaviour called "gen_server", and the default timeout is 5000 milliseconds. If it timeouts an exception is raised, and a common strategy is to simply let the process crash, and restart to try again [0].
On the other hand it's not async call all the way down: it's only the calls between actors that are asynchronous: if you write a loop [1] within your actor then it's synchronous. To write asynchronous code you need to spawn new actors (which is trivial, and will always by default result in an exception being raised if something get lost or timeout after 5 seconds).
Note that when you write Erlang code the granularity of an actor is absolutely not the same as when you write OOP code: you don't wrap every data structure in a process, rather processes tend to be about processing data, or abstracting an IO point (a file, a socket, etc.) since the semantic is perfect.
[0] The implementation of all of this takes more or less 1 line of code, this is really the defaults.
[1] you need to write it recursively in Erlang.
Erlang == "independent reliable units of computing". (I think Erlang programmers call them "functions"? Or maybe "processes"? AFAIR only the "process" level is subject to message loss, yes?)
I'm sure you're right about most of what you wrote, but that's not the point I was getting at.
EDIT: Just to expound: In Kay's conception even e.g. "==" inside a function would be a method invocation and it would be asynchronous (presumably taking two continuations, one for "equal" and another for "not equal"). This absurdity is why I fundamentally disagree with the Actor model (as a model of computation), but also with extremism such as Kay's.
EDIT: Btw, I call bullshit on the 99.9999999% reliable service unless you can come up with a reference. I know that Erlang/OPT is known for legendary relaibility, but I'm unsure on what that means: Does it mean "no interruption" on any given stream (e.g. phone call). Does it mean "no more than X connections fail"?
It's interesting that detailed implementations of the idea become people's definitions (i.e. taking method-based dispatch as the analog to intercellular interaction). With Erlang, the method dispatch becomes modal via the currently active (or lack thereof) receive block in any given process, and I wonder if this captures something too specific or if this difference is worth making at the level of Alan Kay's OOP.
With that in mind, I think you're overstating the orthodoxy. The symbolic representation and execution semantics of messaging models aren't exactly the same. There is no reason == needs to be a "method invocation". If we step back, there isn't really a good definition of what it really means in todays runtimes and compilers anyway. Things get inlined, fast paths get generated for common cases, &c.. His point was more that late binding allows the system to avoid upfront knowledge or hard coding of all useful configurations in a piece of software. Leaving optimization to compilers and runtime systems seems to be the status quo these days.
EDIT: Just on a minor point:
> There is no reason == needs to be a "method invocation".
Well, except that's what the Actor model and, indeed, Smalltalk (Kay et al.'s project) wants us to believe. I, of course, agree completely with you on this point. It's just a function -- simple as that.[2]
[1] Yeah, yeah, I'm exaggerating, but he did kind of bring on himself by continuing on his quest :). I bear no particular ill will towards him, but I do think he's caught in a web of having-to-be-brilliant-based-on-prior-performance-but-does-not-have-much-to-contribute... as befalls many academics and 'visionaries' (e.g. Uncle Bob with his recent rants). (Disclosure: I used to be an academic... and dislike Unclue Bob because he's far more dogmatic than he deserves to be based on documented experience.)
[2] Weeeeelllll, except to be able to decide equality you need to break the abstraction barrier. Which leaves an abstraction-committed OOP programmer with a bit of a dilemma. Me... I just choose the "Eq" type class. Problem solved and I didn't even have to expose any internals. Win win!
I think the whole pedantic "Alan Kay's OOP" thing comes from his coinage of the term vs. the practical application stemming from the Simula school of object orientation. I have no horse in the race so I really wish one sense of the term would be renamed.
That is the point behind object/actors/processes all the way down. As you only talk through message passing, noone has to know if you are sync or not inside the actor. You have a contract.
Moving a local computation to a remote server often fails because clients have implicit assumptions about performance, but nobody is explicit about deadlines except in hard real-time programming. So in practice you get timeouts at best.
You should check out the video[0] of Joe Armstrong interviewing Alan Kay - it seemed to me that Erlang was in some respects closer to what Alan Kay intended for OOP than even smalltalk was.
Regarding the 99.9999999% reliability, it's from Joe Armstrong about a system called AXD301. Here is the quote [0]: "Erlang is used all over the world in high-tech projects where reliability counts. The Erlang flagship project (built by Ericsson, the Swedish telecom company) is the AXD301. This has over 2 million lines of Erlang. The AXD301 has achieved a NINE nines reliability (yes, you read that right, 99.9999999%). Let’s put this in context: 5 nines is reckoned to be good (5.2 minutes of downtime/year). 7 nines almost unachievable ... but we did 9."
about the nine nines : it was a british telecom company that claimed it, by extrapolating from a 14 node network during 8 months. I prefer to say that erlang make it easier to reach five nines, then enable you to go to seven nines if you really need it.
The actor model is a great way to unify remote and local processes with identical syntax, and solves a ton of syntax related issues such as overriding the default '==' operator, etc.
OOP more naturally models a multicore system, because it bundles state and behavior. That more closely modelsthe hardware: multiple CPUs, each of which address a non-uniform portion of memory. That is like methods an an object that operate on local/private state.
What's missing in some languages is an abstraction for "thread of control", i.e. your program counter.
But that can be added, and has been -- people are just not using it as widely as they should. Whereas I think functional languages are great for some things but awkward for others. I like algebraic data types but I'm beginning to think they are orthogonal to functional programming, as in Rust, which has them, but is an imperative language.
I expanded on this style of treating OO languages in a more functional style state here:
Hadoop is written in an OO language -- Java. But it has a functional architecture (MapReduce).
So really the two camps aren't that far apart. They both have something to bring to the table.
Opponents of OOP are thinking of all the spaghetti imperative code with globals dressed up as classes, which people have realized is a mistake, but there's still hundreds of millions of lines of code out there like that. And there are tons of "frameworks" which actually FORCE you to write in that style.
But yeah there is no real contradiction that I see. Thinking that you can write a for loop to update a hash table is as big a mistake as writing spaghetti objects.
(If you don't do lenses or similar, it's still pretty tedious to deal with "deep updates", but that's coincidental. Lenses do solve this problem, although they're awful in practice in Scala.)
Totally agree if you're talking about e.g. Java or Python or even C#.
I don't think this is a problem ingrained in OO. You could very easily conceive of an "OO" language that has separate constructs for objects (mutable) and data (immutable) and there would be no problem at all. Immutable data just happens to be common in functional languages but there's no reason you couldn't introduce it to an "OO" language.
It is also functional with immutable data and variables. So it requires a different mindset, but I find it it fits well with the actor model. But it certainly is not coupled to it in general.
Scala is an example of a pure OO language with immutable data structures.