Could anybody elaborate on this? How does an actor differ from an object that uses promises to talk to a server?
Could anybody elaborate on this? How does an actor differ from an object that uses promises to talk to a server?
Actors are like objects, but the only way to interact with them is sending them a message - you can’t directly access their properties, call their methods, etc. Messages are bits of immutable data sent between actors asynchronously, but those messages go into the actor’s mailbox, and then the actor processes them synchronously, one at a time. You can get parallelism by adding more actors of the same type (10 actors of the same type lets you process up to 10 messages concurrently). Because actors are totally synchronous INTERNALLY, and nobody can directly access their internal state, it’s fine for them to have mutable internal state, with no locking/synchronization needed, and much easier to reason about than mutable state normally is in highly concurrent applications.
With that being said, the actor model has a lot of inherent complexity. If you don’t need it, don’t use it. But if you really do need lots of concurrency and lots of mutable state, it’s often a great choice.
They’re also good for distributed systems. Actors only communicate by sending messages to other actors, but it doesn’t really matter if that message is sent in memory to an actor in the same process, or over the network to an actor on a different machine. So you can take a single actor system and split it across many machines pretty easily. Generally the actor framework you’re using will handle delivering messages over the network, ensuring there’s the right number of actors running across all nodes combined, etc.
While the implementation in many popular, e.g. C++ and family descended from it, languages muddies the waters, basically methods in OO are a convenient way to define handlers for messages sent to objects, and in pure OO you also can't interact with objects except by sending messages to them. The difference between the basic OO model and the actor model is that OO message sends are synchronous request/response and actor model message sends are asynchronous fire-and-forget.
Sure, you can do async request/response in the actor model, but the low-level basic mechanism all communication is built on is send a message to an actor mailbox and keep going. Everything else is built in top of that.
Yes, Goblins' "vat" model descends from E, which supports the "hybrid worldview". http://www.erights.org/
I have a tendency to call Goblins objects "actors", though some people in the community like to point out that "distributed object" is preferable since synchronous call/return is not supported in actors. But yeah.
It doesn't need to be distributed, and it's better not to include that because the approach to error handling gets a lot easier. If you're distributed every call can fail, be randomly slow, has marshaling costs, is untrusted, etc.
Kubelet wasn't necessarily designed as an actor. I think it's a concrete thing that people have interacted with that is a pretty good example of an actor. It has it's own control loop. It sits there and actively monitors the state of the node. It constantly reports the state. If something's wrong, it can try to correct it.
Most services and objects just sort of sit around and wait to be called. Kubelet is a little different, it's got a main driving thread that's constantly looking for trouble, and acting on what it finds.
The line is still pretty fuzzy for me, but maybe this helps someone connect some conceptual dots, to distinguish between a service and an actor.
Not trying to be pedantic, but that's technically not true of the Actor formalism [0]. Quoting wikipedia:
> an actor can designate the behavior to be used to process the next message, and then in fact begin processing another message M2 before it has finished processing M1.
Available implementations (such as Erlang, Akka, Pony) do not support that pipelining so are single threaded per actor as you state though. As an aside, the behaviour you describe is part of the CSP formalism [1] - a related but different approach to concurrent systems.
[0]: https://en.wikipedia.org/wiki/Actor_model#Inherently_concurr...
[1]: https://en.wikipedia.org/wiki/Communicating_sequential_proce...
Perhaps some of the blame belongs on the creators of Erlang for kind of jumping on the "actor" (and later, "OOP") bandwagons as marketing? to try to make Erlang less scary.
> An actor is a computational entity that, in response to a message it receives, can concurrently:
send a finite number of messages to other actors;
create a finite number of new actors;
designate the behavior to be used for the next message it receives.
> There is no assumed sequence to the above actions and they could be carried out in parallel.All above are true of Erlang except intra-actor concurrency. Even the last bullet on designating behaviour: An Erlang process (actor equivalent) can decide to use a different function when receiving the next message. It just can't pipeline handling messages.
So Erlang isn't that far removed from the Actor model in its behaviour, even if it wasn't designed as an Actor system in the first place. One might say actors are an (approximate) emergent property of Erlang rather than an intentionally designed one.
Slightly off-topic but the only system I've used that was designed & built from the outset as an actor system is Rosette [1]. It was a really interesting project, and does support pipelining, but has been dormant for more than a decade.
[0]: https://en.wikipedia.org/wiki/Actor_model#Fundamental_concep...
[1]: https://github.com/leithaus/Rosette
--
EDIT: clarified that Rosette is the only system I've used that was explicitly based on the Actor model from the outset.
Point is all of these deviations from actor system were choices made by the Erlang team in the name of pragmatism. They all exist because there was a use case and the first teams using Erlang needed them for something real; that Erlang is not an actor system is important, because it's a highly pragmatic system -- not one that is based in theory.
Ets tables are observationally equivalent to an Erlang process per table and sending it a message for each function call. It's not actually implemented that way, but I don't think that is grounds to disqualify it. I don't remember enough details about the naming process, but I think that might be similar; you could send a message to a naming process to set and lookup the names (although you'd have a bit of trouble finding out what the process id of the naming process is, wouldn't you?), it's just not very pragmatic.
Nifs certainly have the potential to break the model of course. I'd think selective receive should be fine too, it's equivalent to reading (or peeking) and saving messages until a matching message is received, processing that message, then processing future messages from the saved queue. It's just implemented in a more pragmatic way.
If immutability is important, then the process dictionary means Erlang doesn't qualify, but again, it's pragmatic.
I'd rather have a pragmatic almost actor system than a dogmatic Actor system that's hard to use.
I think that is the key insight of actor-systems.
Actors don't have state. But they can calculate their replacement based on their immutable data. Thus the way an actor at a given address evolves is described by the functions that calculate the successor actors.
Thus you get Pure Functional Programming implemented on top of a fabric of distributed evolving entities. You can understand the behavior and evolution of such a system as a composition of function-calls, where functions always produce the same result for the same arguments.
Perhaps we have different ideas of what an "actor system" is though.
Same question applies to Erlang, which I realize is not exactly an actor system as per the below comment, and has a much more sophisticated error recovery story with supervisors, but the general question holds.
If your services don't deal with internal mutable state, nor high degrees of concurrency then there isn't much gain to be had with an actor system. That said, that begs the question of what the queue is for; just create more instances, since there's no internal state to share.
As soon as you start having internal mutable state and high levels of concurrency, that's where the actor model applies. Queues don't exist for concurrency (you don't need them; just create more executors), they exist for imposing sequence where it is needed (an obvious case; you have a DB connection, you want to only have one query at a time. So every desired query goes into a queue, and the process at the end that owns the DB connection pulls from it). Internal mutable state gets stored inside of an actor; updates and reads get serialized on that actor.
At the highest level, I would describe the actor model as taking a 'successful' model for distributed computing, and making it the only model you use, even locally.
In, let's say Java, for instance, using standard concurrency approaches, it matters where a process lives. My way of operating/communicating to another thread of execution (unit of concurrency) is very, very different than my way of operating/communicating to another machine. Locally I have threads and locks and need to be very mindful. When communicating to another machine, I send a message and that's it (maybe I expect a response, and timeout if I don't get one, but that's really just the same thing, the other machine sending a message).
I don't actually need a queue involved for theoretical correctness unless I need to process messages in sequence (after all, I could have multiple copies of the other process, and send a message to each of them). Now, in the real world I do, simply because if my concurrency gets too large it can't be handled by what the units of concurrency already available (instances, threads, whatever), and scale up takes time, but that's really just a special case of why I need to impose a sequence on messages (handle these first, then handle these, rather than handle all of them concurrently).
The actor model makes this the local model of communication (and so makes the impedance mismatch negligible between local and distributed; so much so that some languages it's actually irrelevant whether you're sending a message to a local actor, or a remote one). Scaling concurrency up internally just means spin up a new actor. When you need to serialize, you send messages to the same actor, where it ends up in a queue.
So it's not solving a different problem, exactly, if that problem is "how do we write systems that can do multiple things at once", but the specifics, complexity, etc, tend to be pretty different. The problems it's solving are a bit more subjective than simply "can we handle this problem", and more "how well do our tools and mental model lend themselves to the problem we're trying to solve".
I kind of disagree with this. I mean, it's kind of correct, if you literally have a sequential problem with immutable state you're gaining nothing from it, but you're also losing nothing (you...will have one actor, no messages being passed, so the only cost is the syntax of the language; the actor model isn't adding any complexity because you aren't using it).
But, a lot of problems we've historically learned to view as sequential are in fact concurrent. Deeply concurrent. We've created entire concepts specifically to impose sequential processing upon innately concurrent activities.
As an example, almost everywhere we use queues we could instead model as unrelated processes able to be run concurrently. You still may have a bit of queueing/scheduling for unbounded processes, but I know from production experience that that is far, far simpler to do and get right using actors than a traditional queue/priority queue and worker(s).
And that's where it shines. It encourages you to start thinking about what in fact -should- be done concurrently, by making that easy rather than a chore as it is in most other languages, and that leads to -less- complexity.
Is it the supporting structure surrounding an actor focused library/language? Lots of person-hours have gone into making them to work well. Some home grown queue + pool of threads working off that queue... not so much.
Some examples:
- Some common patterns. E.g. actor hierachy and supervision, circuit breaker.
- Persistence. The actor persists its state (changes) and can wake up from a sleep.
- Clustering. The actor can live in another node, and you interact with it using the same API.
What I meant is that you can launch a co/go-routine and have a channel that only it can read. That is like 70% the power of having an actor system.
---
> an actor vs a queue with workers
See the sibling comment.
TLDR: Actors each has their own queue. The orderings in each queue are independent. This can help with parallelization if it fits the shape of the problem.
You can also iterate over an array in parallel and only modify the current element and that is like 70% of the power of having an actor system.
"Structured concurrency is a programming paradigm aimed at improving the clarity, quality, and development time of a computer program by using a structured approach to concurrent programming." - Wikipedia
Now, using actors, this was trivial to do. Each actor was basically a glorified state machine, with each command, status update, etc, a message coming in that updated the internal state, and caused it to change state, and, where appropriate, update anything it needed to (such as the internal timer for when it would stop the job). We loaded up everything that had to happen in the next 12 hours, via an actor that would wake up and load more periodically (with an actor registry to find and confirm what existed already, as well as used to route incoming commands from users to change jobs). There is no complexity in managing a global state of priority between the actors; that's a 'solved problem' by the time slicing the underlying actor engine is doing for us; it's well tested, well proven, and invisible to us. Our implementation allows us to treat each job as existing in isolation, and write our code as though it and it alone is the job that needs to be executed. There is no global synchronization state to manage (well, with the caveat of taking the infinite list of future jobs, and ensuring we have actual processes equating to those scheduled to start in the next 12 hours, but that's a comparatively simple problem, and a proper use of sequential ordering. We want to execute on these jobs first), nor should their be based on the problem we're trying to solve.
Using a queue and traditional threading model though? Well, we'd need a synchronized priority queue, as items could have their priority changed at any time, from multiple sources. We'd also still need a registry. We'd need a pool of workers, that are mostly just blocking...hope we don't have > (number of workers) things needing to execute concurrently. Our complexity to manage all of this is high; the level of testing needed to validate that there aren't emergent bugs, also high. Our semi-realtime requirements are probably out the window (high contention on the queue, or high enough concurrency to exhaust the thread pool could lead to very measurable delays in updates). And ultimately, it all comes to the fact that the queue exists to pretend that the state of the world is sequential, and that we then had to then contort ourselves around to bring back to a state that allows for suitable concurrency. Updates to a job can end up killing a worker and inserting something back onto the queue, can remove something from the queue and reinsert it at a different position, etc, all while other workers are trying to access the queue simultaneously.
Actors are the simpler solution.
The comment about an actor's mailbox being a queue is missing the point (though, amusingly, demonstrating it). Queues are sequential. Actors are capable of doing only one thing at a time, and so when they are requested to do multiple things, they will do them sequentially. A mailbox filling with things that don't demand sequential processing is a code smell in an actor system. This is useful if you have limited resource access (queries on a DB connection, say), or need things to be done sequentially, and are very much akin to threads. But if you actually want multiple things concurrently...you should have multiple actors, not send them into the mailbox of one actor (which is very much the queue situation mentioned above; you imposed a sequential ordering on things you want done concurrently. Why?). That's the key difference; in most other languages, the queue is taking multiple things you actually want concurrently, and imposing a sequence to them. In the actor model, your queue is in fact things you DO want done sequentially; the things you want done concurrently should get federated out to multiple actors.
In short: different computational model -- Actor model vs. Communicating sequential processes.
This really helped me: https://youtu.be/7erJ1DV_Tlo
Edit: added link to video
However, while the single-threadedness is a hard requirement and behavior of the system, the rule that only one message is processed at a time is not.
In Orbit (and Orleans) the ability to process multiple messages at the same time is called reentrancy. it can be specified at the actor or method level usually.
So in your example, say your initial message kicks off a write that is going to take 30 seconds. You start the write and assuming it's async, the actor then suspends while it waits for that to complete. While the actor is suspended another message asking for the progress arrives and that message is reentrant, it can process that message and respond with the progress, then the actor suspends again and at some point another reentrant message arrives or the initial write finishes and the first message finally gets a response. If a none reentrant message is received it will queue in the mailbox until after the first one is completed.
So, with reentrancy you still get the single threaded guarantee but you effectively get interleaving of messages when the actor is suspended waiting for some other async task to complete. This means that if you have an actor with reentrant messages you have to make that code safe to run during any potential interleaving point in another message (for example an await in async C#).
Hopefully that made sense, Orleans has some docs about reentrancy here you might find useful: https://dotnet.github.io/orleans/docs/grains/reentrancy.html
Other actor frameworks (particularly ones that are not based on virtual actors) may have different guarantees and behaviors but from what I've seen there is always some equivalent.
Then what's a virtual actor and how does it differ? You'll eventually end up at the Orleans paper from MSFT research:
https://www.microsoft.com/en-us/research/publication/orleans...
It really has nothing to do with promises (as used colloquially) and is a way to design large distributed systems using message passing and localized state that's persisted.
1. Communication solely by asynchronous message passing
2. A mailbox per actor that means the reception of messages and the processing of them is decoupled (i.e. an actor can receive new messages even when processing a previously-received message).
The OO model in general supports that paradigm. Pretty much all mainstream OO languages are synchronous by default though. Call a method and the caller is suspended until the method returns. Multiple clients can call methods on the same object concurrently, but doing that safely requires some form of locking/protection to be implemented in the target object. Stated alternatively: threads in mainstream OO languages run "across" objects, whereas each actor has its own thread in the Actor model.
There are certainly similarities - in the sense that both actors and objects encapsulate state, with reading/writing occurring through well-defined interfaces (messages and methods respectively). But the threading model is quite different - it's not just a matter of scale.
[0] https://en.wikipedia.org/wiki/Actor_model
EDIT: corrected grammar & formatting.
But I’m actor systems the actors maybe more stateful that you typical object and be location transparent. Communication between actors does require knowledge of if another actor is in process or remote.
Also the actor supervising model is a completely different way of recovering from errors than in OO systems.
When people talk about actor-based languages or frameworks however, theyre suggesting syntax primitives which help express this more naturally.
The simple pitch is: OO where objects could be on any machine.